GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.AI 2502.07266

When More is Less: Understanding Chain-of-Thought Length in LLMs

Study finds LLM reasoning accuracy follows an inverted U-shaped curve with CoT length; proposes length optimization methods.

Yuyang Wu, Yifei Wang, Ziyu Ye et al.

2025-02-11 24
cs.AI 2501.17811

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Janus-Pro scales up to 7B parameters, employing optimized training, expanded data, and decoupled visual encoding, achieving state-of-the-art multimodal understanding and text-to-image generation.

Xiaokang Chen, Zhiyu Wu, Xingchao Liu et al.

2025-01-30 841 citations 40
cs.AI 2501.17161

SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

This study compares SFT and RL in foundation models, showing RL achieves 77.8% success in OOD generalization, outperforming SFT by 33.8%.

Tianzhe Chu, Yuexiang Zhai, Jihan Yang et al.

2025-01-29 26
cs.AI 2501.15147

A Causality-aware Paradigm for Evaluating Creativity of Multimodal Large Language Models

Introduced LoTbench framework to evaluate creativity of multimodal LLMs using Oogiri game, finding the gap with humans is small.

Zhongzhan Huang, Shanshan Zhong, Pan Zhou et al.

2025-01-25 0
cs.AI 2501.12599

Kimi k1.5: Scaling Reinforcement Learning with LLMs

Kimi k1.5 scales LLMs with RL, achieving 77.5 on AIME and other top scores.

Kimi Team, Angang Du, Bofei Gao et al.

2025-01-22 27
cs.AI 2501.09685

Inference-Time Alignment in Diffusion Models with Reward-Guided Generation: Tutorial and Review

Unified inference-time guidance framework for diffusion models using reward and value functions, improving protein design performance without model fine-tuning.

Masatoshi Uehara, Yulai Zhao, Chenyu Wang et al.

2025-01-17 76 citations 40
cs.AI 2501.09136

Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG

Agentic RAG enhances real-time response by embedding autonomous AI agents for dynamic retrieval and generation.

Aditi Singh, Abul Ehtesham, Saket Kumar et al.

2025-01-16 2
cs.AI 2501.05366

Search-o1: Agentic Search-Enhanced Large Reasoning Models

Search-o1 integrates agentic retrieval and document reasoning, boosting large reasoning models' performance on complex tasks.

Xiaoxi Li, Guanting Dong, Jiajie Jin et al.

2025-01-10 43
cs.AI 2501.04410

User Simulation in the Era of Generative AI: User Modeling, Synthetic Data Generation, and System Evaluation

Generative AI-based user simulation using large language models enhances behavior modeling, synthetic data generation, and system evaluation, advancing personalization and safety.

Krisztian Balog, ChengXiang Zhai

2025-01-08 35
cs.AI 2501.02497

A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning

This survey links Tent, CoT, PRMs, and MCTS to explain how test-time compute moves models from intuition toward deliberate reasoning.

Yixin Ji, Juntao Li, Yang Xiang et al.

2025-01-05 17
cs.AI 2412.19723

OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis

OS-Genesis employs reverse task synthesis via interaction-driven exploration, significantly improving GUI trajectory quality and diversity, boosting model performance by nearly 80% on benchmarks.

Qiushi Sun, Kanzhi Cheng, Zichen Ding et al.

2024-12-28 141 citations 27
cs.AI 2412.18985

TravelAgent: Generative Agents in the Built Environment

TravelAgent combines generative agents with 3D environments, achieving 76% task completion over 1898 steps, enhancing urban behavior simulation.

Ariel Noyman, Kai Hu, Kent Larson

2024-12-26 35
cs.AI 2412.18424

LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating

Introduces LongDocURL, a comprehensive multimodal benchmark for long document understanding, with 20 sub-tasks, 2,325 high-quality QA pairs, revealing significant performance gaps.

Chao Deng, Jiale Yuan, Pi Bu et al.

2024-12-24 86 citations 38
cs.AI 2412.17287

LLM4AD: A Platform for Algorithm Design with Large Language Model

LLM4AD unifies LLM-driven algorithm search; EoH, FunSearch, and (1+1)-EPS beat random sampling on most of nine tasks.

Fei Liu, Rui Zhang, Zhuoliang Xie et al.

2024-12-23 14
cs.AI 2412.09413

Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Proposes STILL-2 framework combining imitation, exploration, and self-improvement to develop industry-level slow-thinking reasoning systems.

Yingqian Min, Zhipeng Chen, Jinhao Jiang et al.

2024-12-13 39
cs.AI 2412.06771

Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty

Proactive T2I agent uses belief graphs and multi-turn questioning to improve alignment, achieving 2x VQAScore over standard methods.

Meera Hahn, Wenjun Zeng, Nithish Kannen et al.

2024-12-10 39
cs.AI 2412.05255

TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft

TeamCraft builds a Minecraft-based benchmark with 55,000 multimodal multi-agent tasks to evaluate generalization in complex environments.

Qian Long, Zhi Li, Ran Gong et al.

2024-12-07 36
cs.AI 2412.05167

Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models

Proposed ADU-Bench evaluates 16 LALMs across multi-scenario, multi-skill, multi-language, and ambiguity tasks, revealing performance gaps and strengths.

Kuofeng Gao, Shu-Tao Xia, Ke Xu et al.

2024-12-07 32
cs.AI 2412.04782

A Survey of Sustainability in Large Language Models: Applications, Economics, and Challenges

Integrating energy-efficient techniques and lifecycle assessment reduces LLMs' environmental impact by over 50%.

Aditi Singh, Nirmal Prakashbhai Patel, Abul Ehtesham et al.

2024-12-06 37
cs.AI 2412.04759

REGENT: A Retrieval-Augmented Generalist Agent That Can Act In-Context in New Environments

REGENT employs retrieval-augmented transformer policies for zero-shot generalization, with 3x fewer parameters and 10x less data, outperforming SOTA in unseen environments.

Kaustubh Sridhar, Souradeep Dutta, Dinesh Jayaraman et al.

2024-12-06 31
Prev 1 ... 32 33 34 35 36 37 38 ... 44 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home