Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
Mobius-v0 separates knowledge and reasoning, achieving nearly 4x inference speedup while maintaining accuracy.
Kai Chen, Jifeng Ding, Ning Ding et al.
Mobius-v0 separates knowledge and reasoning, achieving nearly 4x inference speedup while maintaining accuracy.
Kai Chen, Jifeng Ding, Ning Ding et al.
HELIX employs model-harness co-evolution with source-traceable interventions, boosting task coverage by 4.0% and verified coverage by 58.0% in code repair tasks.
Tianyu Fan, Chao Huang
OmniScientist employs multimodal perception and a multi-agent architecture to automate multidisciplinary research from raw evidence to publication, achieving an average score of 6.3.
Bobo Li, Hao Fei, Tianjie Ju et al.
Proposed a framework to evaluate long-horizon AI research, analyzing 7 models across 36 tasks.
Yiwei Li, Wanli Yang, Hexiang Tan et al.
Uniform Herding dynamically refreshes exemplars in feature space, achieving 44.00% accuracy on CIFAR-100, outperforming iCaRL.
Krishna Subedi
SkillMisevo-Gym reveals skill misevolution in LLMs; SafeEvolve reduces unsafe retrieval by 26.7 percentage points.
Xutao Mao, Liangjie Zhao, Xiang Zheng et al.
This study introduces an automated KG-DML construction framework using RAG and LLMs for complex system diagnostics.
Saman Marandi, Yu-Shu Hu, Mohammad Modarres
This paper introduces budget-dependent evaluation of LLMs, revealing model rank reversals across token budgets (64-4096 tokens), with 14.1% of oracle gap captured by a budget-aware router.
Rodrigo Guedes de Souza, Alison R. Panisson
MindMemOS offers a portable, self-evolving memory layer, achieving 94.03% accuracy on LOCOMO.
Kaichao Liang, Yuqi Cui, Hao Kong et al.
XBRIDGE combines lexical anchor mapping and latent enrichment bridge for efficient heterogeneous LLM communication, reducing latency by 11×.
Wooseong Yang, Wei-Chieh Huang, Weizhi Zhang et al.
Introducing Social Chain of Thought (SCoT), a multi-agent framework that improves differential diagnosis recall by 4-12% through multi-round specialist interactions.
Del Coburn, Scott Sanner, Dan Silver
Using AI to tighten bounds on the Grothendieck constant, achieving a lower bound of 6π/11≈1.7135, surpassing previous 1.6769.
Alan Li, Rahul Saha, Anton Xue et al.
SkillLens uses Visual Skill Cards to improve GUI action prediction, boosting Step SR+11.6.
Zhou Liu, Ligang Huang, Zeli Su et al.
Proposes a mixed-state quantum prototype framework for incremental learning, addressing capacity limits with minimal qubits and robust distance metrics.
Yu Wu, Qianli Zhou, Xinyang Deng et al.
Introduces DSAgentBench, a benchmark for evaluating AI agents' ability to automate full data science workflows in real OS environments, achieving only 56.7% success.
Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub et al.
MESA selects complementary memory structures per query, reaching 65.1% on AMA-Bench with 41% fewer evidence tokens.
Beidi Zhao, Yaoqi Chen, Yuru Feng et al.
TIDE method corrects teacher-student mismatch via bounded Hellinger shaping and top-K injection, boosting Avg@8 from 6.9% to 20.3%.
Zichao Yu, Chengzhi Yu, Shengze Xu et al.
Proposes BCSD, a dual-view self-distillation framework, improving external skill utilization in LLMs; achieves state-of-the-art results on ALFWorld and WebShop.
Tianjun Pan, Yuan Li, Hongda Wang et al.
Coderlet employs a request lifecycle-driven architecture, explicitly separating model, execution, and state boundaries for continuous AI interaction.
Mengfan Li
CoRE uses graph-based dominant set extraction and replicator dynamics to improve test-time RL rewards, boosting accuracy by 21.7 points.
Ambuj Mehrish, Sebastiano Vascon