More Than Can Be Said: A Benchmark and Framework for Pre-Question Scientific Ideation
InciteResearch framework transforms tacit understanding into explicit research proposals, enhancing novelty and impact.
Jie Yu, Song Qiu
InciteResearch framework transforms tacit understanding into explicit research proposals, enhancing novelty and impact.
Jie Yu, Song Qiu
SPEED method reduces long-context inference cost via shallow prefill and deep decode, enhancing efficiency.
Jungsuk Oh, Hyeseo Jeon, Hyunjune Ji et al.
VibeServe employs multi-agent loops to automatically generate bespoke LLM serving systems, matching or exceeding hand-tuned performance.
Keisuke Kamahori, Shihang Li, Simon Peter et al.
SkillRet is a large-scale long-text skill retrieval benchmark with 17,810 skills, using structured semantic tags to improve LLM retrieval performance.
Hongcheol Cho, Ryangkyung Kang, Youngeun Kim
Introduces knowledge-graph paths as intermediate supervision, improving self-evolving search agents' QA accuracy by 7.3% on multi-hop tasks.
Huyu Wu, Jun Liu, Xiaochi Wei et al.
Using an information-theoretic and geometric framework, the paper reveals that attention mainly reconfigures representations while FFN drives semantic innovation, exposing redundancy in LVLMs.
Gongli Xi, Ye Tian, Mengyu Yang et al.
SPARK uses knowledge graphs for structured multi-hop reasoning, outperforming unstructured baselines with 93% accuracy on ScienceQA.
Hyobin Park, Taeseop Kim, Dong-Geol Choi
Proposes SCPRM, integrating schema-aware distance for risk-sensitive multi-hop KG reasoning, improving accuracy by 1.18%.
Jiujiu Chen, Yazheng Liu, Sihong Xie et al.
Foundation-model-based industrial agents leverage large language models for autonomous decision-making, enhancing human interaction (+37%) and uncertainty handling (+35%).
Vincent Henkel, Felix Gehlhoff, David Kube et al.
MILD integrates bidirectional perception and multi-layered alignment with ECPO optimization for human-vehicle collaboration.
Jiyao Wang, Yunbiao Wang, Yubo Jiao et al.
DiagramNet employs multi-stage training and a decoupled multi-agent workflow to achieve superior system-level diagram recognition, outperforming GPT-5 and industry benchmarks.
Jincheng Lou, Ruohan Xu, Jiapeng Li et al.
GUI-SD introduces on-policy self-distillation with visual privileged context and entropy-guided loss, boosting GUI grounding accuracy to 68.4% and training speed 4× faster.
Yan Zhang, Daiqing Wu, Huawen Shen et al.
This paper introduces a reinforcement learning framework for GUI agents, categorized into offline, online, and hybrid strategies, emphasizing reward engineering and data efficiency.
Junan Hu, Jian Liu, Jingxiang Lai et al.
Proposes an agentic RL framework for LLMs integrating meta-reasoning, multi-step planning, and external tools, achieving 50% improvement on complex tasks.
Fangming Cui, Ruixiao Zhu, Cheng Fang et al.
Proposed event-driven step-level cascade framework improves efficiency, achieving 58.2% success rate.
Jinbiao Wei, Kangqi Ni, Yilun Zhao et al.
Proposed a disagreement-guided strategy routing method, improving mathematical reasoning accuracy by 3%-7%.
Zhimin Lin, Yixin Ji, Jinpeng Li et al.
SciCrafter evaluates AI's discovery-to-application ability in Minecraft; current models achieve only 26% success.
Zhou Ziheng, Huacong Tang, Jinyuan Zhang et al.
ZenBrain is a 7-layer neuroscience-inspired memory architecture integrating 15 mechanisms, significantly improving QA and cross-session reasoning with 91.3% oracle accuracy at 1/106 query cost.
Alexander Bering
Proposes an LLM-based evaluation framework to enhance math reasoning assessment accuracy beyond symbolic math limitations.
Erez Yosef, Oron Anschel, Shunit Haviv Hakimi et al.
AgentSearchBench improves agent search ranking quality using execution signals, bridging the gap between semantics and performance.
Bin Wu, Arastun Mammadli, Xiaoyu Zhang et al.