When More is Less: Understanding Chain-of-Thought Length in LLMs
Study finds LLM reasoning accuracy follows an inverted U-shaped curve with CoT length; proposes length optimization methods.
Yuyang Wu, Yifei Wang, Ziyu Ye et al.
Study finds LLM reasoning accuracy follows an inverted U-shaped curve with CoT length; proposes length optimization methods.
Yuyang Wu, Yifei Wang, Ziyu Ye et al.
Janus-Pro scales up to 7B parameters, employing optimized training, expanded data, and decoupled visual encoding, achieving state-of-the-art multimodal understanding and text-to-image generation.
Xiaokang Chen, Zhiyu Wu, Xingchao Liu et al.
This study compares SFT and RL in foundation models, showing RL achieves 77.8% success in OOD generalization, outperforming SFT by 33.8%.
Tianzhe Chu, Yuexiang Zhai, Jihan Yang et al.
Introduced LoTbench framework to evaluate creativity of multimodal LLMs using Oogiri game, finding the gap with humans is small.
Zhongzhan Huang, Shanshan Zhong, Pan Zhou et al.
Kimi k1.5 scales LLMs with RL, achieving 77.5 on AIME and other top scores.
Kimi Team, Angang Du, Bofei Gao et al.
Unified inference-time guidance framework for diffusion models using reward and value functions, improving protein design performance without model fine-tuning.
Masatoshi Uehara, Yulai Zhao, Chenyu Wang et al.
Agentic RAG enhances real-time response by embedding autonomous AI agents for dynamic retrieval and generation.
Aditi Singh, Abul Ehtesham, Saket Kumar et al.
Search-o1 integrates agentic retrieval and document reasoning, boosting large reasoning models' performance on complex tasks.
Xiaoxi Li, Guanting Dong, Jiajie Jin et al.
Generative AI-based user simulation using large language models enhances behavior modeling, synthetic data generation, and system evaluation, advancing personalization and safety.
Krisztian Balog, ChengXiang Zhai
This survey links Tent, CoT, PRMs, and MCTS to explain how test-time compute moves models from intuition toward deliberate reasoning.
Yixin Ji, Juntao Li, Yang Xiang et al.
OS-Genesis employs reverse task synthesis via interaction-driven exploration, significantly improving GUI trajectory quality and diversity, boosting model performance by nearly 80% on benchmarks.
Qiushi Sun, Kanzhi Cheng, Zichen Ding et al.
TravelAgent combines generative agents with 3D environments, achieving 76% task completion over 1898 steps, enhancing urban behavior simulation.
Ariel Noyman, Kai Hu, Kent Larson
Introduces LongDocURL, a comprehensive multimodal benchmark for long document understanding, with 20 sub-tasks, 2,325 high-quality QA pairs, revealing significant performance gaps.
Chao Deng, Jiale Yuan, Pi Bu et al.
LLM4AD unifies LLM-driven algorithm search; EoH, FunSearch, and (1+1)-EPS beat random sampling on most of nine tasks.
Fei Liu, Rui Zhang, Zhuoliang Xie et al.
Proposes STILL-2 framework combining imitation, exploration, and self-improvement to develop industry-level slow-thinking reasoning systems.
Yingqian Min, Zhipeng Chen, Jinhao Jiang et al.
Proactive T2I agent uses belief graphs and multi-turn questioning to improve alignment, achieving 2x VQAScore over standard methods.
Meera Hahn, Wenjun Zeng, Nithish Kannen et al.
TeamCraft builds a Minecraft-based benchmark with 55,000 multimodal multi-agent tasks to evaluate generalization in complex environments.
Qian Long, Zhi Li, Ran Gong et al.
Proposed ADU-Bench evaluates 16 LALMs across multi-scenario, multi-skill, multi-language, and ambiguity tasks, revealing performance gaps and strengths.
Kuofeng Gao, Shu-Tao Xia, Ke Xu et al.
Integrating energy-efficient techniques and lifecycle assessment reduces LLMs' environmental impact by over 50%.
Aditi Singh, Nirmal Prakashbhai Patel, Abul Ehtesham et al.
REGENT employs retrieval-augmented transformer policies for zero-shot generalization, with 3x fewer parameters and 10x less data, outperforming SOTA in unseen environments.
Kaustubh Sridhar, Souradeep Dutta, Dinesh Jayaraman et al.