LightMem: Lightweight and Efficient Memory-Augmented Generation
LightMem mimics human memory, organizing info into three stages, boosting long-term reasoning with 7.7% accuracy gain and 38× token reduction.
Jizhan Fang, Xinle Deng, Haoming Xu et al.
LightMem mimics human memory, organizing info into three stages, boosting long-term reasoning with 7.7% accuracy gain and 38× token reduction.
Jizhan Fang, Xinle Deng, Haoming Xu et al.
Proposes Enterprise Deep Research (EDR), a multi-agent framework achieving state-of-the-art enterprise analytics with superior performance metrics.
Akshara Prabhakar, Roshan Ram, Zixiang Chen et al.
FARE uses iterative rejection sampling SFT to outperform 70B+ evaluators in reasoning domains.
Austin Xu, Xuan-Phi Nguyen, Yilun Zhou et al.
LLM-driven industry agents automate complex tasks, enhancing productivity.
Yihong Tang, Kehai Chen, Liang Yue et al.
StreamingThinker enables LLMs to think while reading, reducing 80% token wait and over 60% latency, matching batch performance.
Junlong Tong, Yingqi Fan, Anhao Zhao et al.
A knapsack-inspired online framework dynamically tests and optimizes agent component selection, boosting success rates and reducing costs.
Michelle Yuan, Khushbu Pahwa, Shuaichen Chang et al.
Confidence estimation based on activation signals improves LLM trustworthiness, achieving 95% accuracy with reduced latency.
Zhiqi Huang, Vivek Datla, Chenyang Zhu et al.
This study shows that AI models are more persuasive when defending positions aligned with their prior beliefs, using sequential and simultaneous debate protocols, with models favoring sycophantic strategies in conflict scenarios.
María Victoria Carro, Denise Alejandra Mester, Facundo Nieto et al.
Omni-Detective pipeline and Omni-Captioner model enhance multimodal fine-grained perception, outperforming SOTA on key benchmarks.
Ziyang Ma, Ruiyang Xu, Zhenghao Xing et al.
DELTA employs a layer-aware token selection mechanism, reducing attended tokens by up to 4.25× and speeding up inference by 1.54× while maintaining accuracy.
Hossein Entezari Zarch, Lei Gao, Chaoyi Jiang et al.
Proposes SIMBA UQ, a similarity-based framework for black-box uncertainty quantification, improving calibration on QA, summarization, and SQL tasks.
Debarun Bhattacharjya, Balaji Ganesan, Junkyu Lee et al.
Proposes Translation Tangles framework, integrating multi-metric, multi-domain, bias detection for 24 language pairs, analyzing translation quality and biases.
Md. Faiyaz Abdullah Sayeedi, Md. Mahbub Alam, Subhey Sadi Rahman et al.
Proposes PA-Tool, using peakedness to align tool schemas with pretrained models, boosting accuracy by 17% without retraining.
Jonggeun Lee, Woojung Song, Jongwook Han et al.
Customer-R1 employs RL with explicit user personas, boosting next-action prediction accuracy from 7.32% to 39.58%, outperforming baselines.
Ziyi Wang, Yuxuan Lu, Yimeng Zhang et al.
SUPO algorithm scales LLM multi-turn RL via summarization-based context management, enhancing success rate and reducing context length.
Miao Lu, Weiwei Sun, Weihua Du et al.
Analyzed representation flow in SSMs and TBMs using centered kernel alignment and variance metrics, revealing layer-wise information flow differences.
Nhat M. Hoang, Do Xuan Long, Cong-Duy Nguyen et al.
Introduced UserLMs to simulate human behavior in multi-turn dialogues, reducing GPT-4o performance from 74.6% to 57.4%.
Tarek Naous, Philippe Laban, Wei Xu et al.
SimulatorArena uses user profiles to create high-fidelity simulators, achieving ρ=0.7 correlation with human judgments, enabling scalable multi-turn evaluation.
Yao Dou, Michel Galley, Baolin Peng et al.
Even with perfect retrieval, long context length degrades LLM performance; propose shortening context strategy to improve performance.
Yufeng Du, Minyang Tian, Srikanth Ronanki et al.
RE-Searcher combines goal-oriented planning and self-reflection, achieving state-of-the-art robustness and accuracy in complex search environments.
Daocheng Fu, Jianbiao Mei, Licheng Wen et al.