MMToM-QA: Multimodal Theory of Mind Question Answering
Proposes MMToM-QA benchmark integrating Bayesian inverse planning and language models to enhance machine Theory of Mind in multimodal scenarios.
Chuanyang Jin, Yutong Wu, Jing Cao et al.
Proposes MMToM-QA benchmark integrating Bayesian inverse planning and language models to enhance machine Theory of Mind in multimodal scenarios.
Chuanyang Jin, Yutong Wu, Jing Cao et al.
Extends causal models with impossible worlds to formalize mathematical explanations, addressing the issue that all mathematical facts are true in all models.
Joseph Y. Halpern
Proposes NPHardEval, a dynamic benchmark based on complexity classes, with 900 algorithmic questions to evaluate LLM reasoning, updated monthly.
Lizhou Fan, Wenyue Hua, Lingyao Li et al.
Proposes Meta-CPO, integrating differentiable convex programming for safe, rapid adaptation in non-stationary environments.
Minjae Cho, Chuangchuang Sun
Proposed 3DAxiesPrompts (3DAP) embeds 3D coordinate info into images, significantly enhancing GPT-4V's performance in 3D point reconstruction, matching, and object detection.
Dingning Liu, Xiaomeng Dong, Renrui Zhang et al.
Math-Shepherd automatically constructs process supervision data, enabling step-by-step verification and reinforcement of LLMs without human annotations.
Peiyi Wang, Lei Li, Zhihong Shao et al.
SGLang system leverages RadixAttention and compressed FSM to accelerate large language model programs, achieving up to 6.4× throughput increase.
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie et al.
TaskWeaver: a code-first autonomous agent framework supporting rich data structures and dynamic plugin invocation.
Bo Qiao, Liqun Li, Xu Zhang et al.
Proposed AgentMonitor framework achieves 89.4% F1 in safe testing of AutoGPT on the open internet.
Silen Naihin, David Atkinson, Marc Green et al.
Proposes RETROFIT-CQs, leveraging large language models to automatically extract competency questions from existing ontologies, enhancing reusability.
Reham Alharbi, Valentina Tamma, Floriana Grasso et al.
Safety-Gymnasium environment suite and SafePO library enable comprehensive safe RL research with 16 algorithms across diverse tasks.
Jiaming Ji, Borong Zhang, Jiayi Zhou et al.
Empirical analysis of GPT-4’s iterative prompting on graph coloring; reveals limited self-critique and verification abilities, external verifier boosts success to 40%.
Kaya Stechly, Matthew Marquez, Subbarao Kambhampati
MemGPT employs OS-inspired hierarchical memory and function calls to extend LLM context beyond fixed windows, improving long document analysis and multi-session chat.
Charles Packer, Sarah Wooders, Kevin Lin et al.
RoboCLIP leverages pretrained video-language models to generate rewards from a single demonstration, boosting zero-shot robot task success by 2-3x.
Sumedh A Sontakke, Jesse Zhang, Sébastien M. R. Arnold et al.
Goodtriever uses retrieval-augmented methods to reduce inference latency by 43% while maintaining state-of-the-art toxicity mitigation.
Luiza Pozzobon, Beyza Ermis, Patrick Lewis et al.
This study demonstrates that large language models (LLMs) linearly encode truth/falsehood in their internal representations, validated through transfer and causal intervention experiments, especially at scale.
Samuel Marks, Max Tegmark
Corex employs multi-model collaboration (Discuss, Review, Retrieve) to significantly improve complex reasoning, outperforming existing baselines with detailed experimental validation.
Qiushi Sun, Zhangyue Yin, Xiang Li et al.
Framework for LLM-based agents enhances potential for general intelligence.
Zhiheng Xi, Wenxiang Chen, Xin Guo et al.
NExT-GPT achieves any-to-any modality input-output with multimodal adapters and diffusion decoders, tuning only 1% of parameters.
Shengqiong Wu, Hao Fei, Leigang Qu et al.
Confucius framework enhances LLM tool usage via easy-to-difficult curriculum and introspective feedback, surpassing existing baselines.
Shen Gao, Zhengliang Shi, Minghang Zhu et al.