TRWH: A Text-Driven Random Walk Heterogeneous GNN for Semantic-Aware Sparse Recommendation
TRWH combines LLM-generated text features with heterogeneous graphs via random walks, boosting sparse recommendation accuracy.
He Ma, Chen Liu
TRWH combines LLM-generated text features with heterogeneous graphs via random walks, boosting sparse recommendation accuracy.
He Ma, Chen Liu
PreDiff-LM employs hybrid attention, combining pretrained causal transformers with bidirectional denoising, reducing perplexity from 34.1 to 28.7 on WikiText-103.
Zhengtao Yao, Runhao Li, Xupeng Chen et al.
Proposes AUC of Pareto frontier for efficiency evaluation; introduces fluid search, outperforming fixed strategies in 12 tasks.
Haiqian Yang, Yuan Cao
HG-CRC controls hierarchical group risk; on ARC Challenge it achieved 0% empirical violations and WGER=0.
Murilo Salem, Luísa Böhm, Daniel Pontes et al.
SciConsolidate synthesizes procedural knowledge to enhance scientific computing, improving Qwen3.6-27B by 6.26 points.
Liwei Dong, Jiahao Zhao, Nan Xu
RecursiveECG uses evidence-based recursive refinement, transforming expert ECG criteria into validated measurements to improve classifier performance by 10% on benchmark datasets.
Jinliang Deng, Yiming Niu, Yibo Pan et al.
Proposes State Transition Pretraining (STP) using joint inverse and forward dynamics to enhance GUI agents, improving success rates by up to 6.2%.
Xiangyan Liu, Kaixin Li, Haonan Wang et al.
Mechanism mapping reveals evolutionary paths from cognitive architectures to language agents, identifying five residual bundles for future integration.
Haodi Fan, Zucong Lan
CALM trains language models via multi-task RL over structured controllers, enhancing generalization across diverse inference workflows.
Moumita Choudhury, Vanshaj Khattar, Jing Liu et al.
Proposes O²-CritiCuRL, combining offline analysis and online RL to identify critical reasoning steps, boosting multimodal reasoning accuracy and efficiency.
Wendi Deng, Hang Du, Guoshun Nan et al.
TOPOFE, a graph-structured multi-island evolutionary framework, significantly improves AutoFE performance on 29 datasets, surpassing state-of-the-art methods.
Sha Li, Naren Ramakrishnan
Introduced SQBench, a benchmark evaluating language models' task delivery in production workflows, with 220 tasks, combining functional completion and risk assessment, achieving a max of 60.5%.
Summer Sun
SAGE governs AI across its lifecycle; 840 endpoint calls showed low observed harmful compliance under a narrow single-turn protocol.
Mahdi Eslamimehr
HAFS framework improves video quality by 1.5x and reduces response time by 31%.
Goodsol Lee, Juheon Yi, Jinglu Wang et al.
IDEAgent employs a multi-agent Quality-Diversity search framework to generate diverse, high-quality research ideas, outperforming baselines by 3.89× on Yield across 32 topics.
Varun Gumma, Navonil Majumder, Soumitra Sinhahajari et al.
Proposes deployment-feedback-driven continual learning using external memory, boosting τ-bench success rate by 1.6× single-trial, 2.6× with corrections.
Valentin Tablan, Scott Taylor, Kristoffer Bernhem
Introduces MissionBench, a benchmark for zero-shot evaluation of 22 MLLMs on 120 aerial long-horizon tasks, with success rates below 35%.
Suman Navaratnarajah, Taehyoung Kim, Jona Ruthardt et al.
VIGOR uses reward variance to adaptively allocate rollouts, reducing sampling by up to 2.3× while maintaining performance.
Heyang Jiang, Henry Liu, Baharan Mirzasoleiman
The paper defines AI-native systems by revision authority: autonomous implementation rewriting with escalation, verification, and fallback; no empirical metrics are reported.
Cheng Tan
Knowledge-centric self-improvement uses a curated knowledge base to enhance task solving and transferability, reducing costs by 30%.
Xuefei Julie Wang, Lauren Hyoseo Yoon, Chengrui Qu et al.