HiMAP-Travel: Hierarchical Multi-Agent Planning for Long-Horizon Constrained Travel
HiMAP-Travel solves long-horizon travel planning with budget and diversity constraints, achieving 52.65% FPR on the test set.
The Viet Bui, Wenjun Li, Yong Liu
HiMAP-Travel solves long-horizon travel planning with budget and diversity constraints, achieving 52.65% FPR on the test set.
The Viet Bui, Wenjun Li, Yong Liu
DEVS-based framework uses natural language to generate and verify long-horizon discrete-event world models, ensuring consistency.
Zheyu Chen, Huiteng Zhuang, Zhuohuan Li et al.
Distinguishing acoustic and expectation-related ANN representations enhances EEG-based music identification accuracy.
Shogo Noguchi, Taketo Akama, Tai Nakamura et al.
Retrievit combines Transformers and SSMs for efficient in-context retrieval, enhancing data efficiency.
Georgios Pantazopoulos, Malvina Nikandrou, Ioannis Konstas et al.
EvoSkill employs iterative failure analysis to automatically discover and refine skills in multi-agent systems, achieving a 7.3% to 12.1% accuracy boost across benchmarks.
Salaheddin Alzubi, Noah Provenzano, Jaydon Bingham et al.
Conformal Policy Control calibrates likelihood-ratio clipping to explore with finite-sample risk control.
Drew Prinster, Clara Fannjiang, Ji Won Park et al.
Proposed a tensor factorization-based evaluation method combining autorater data and limited human labels for fine-grained generative model assessments.
Felipe Maia Polo, Aida Nematzadeh, Virginia Aglietti et al.
FT-Dojo automates LLM fine-tuning with FT-Agent, excelling in 10 out of 13 tasks.
Qizheng Li, Yifei Zhang, Xiao Yang et al.
Semantic XPath uses tree-structured memory for conversational AI, improving performance by 176.7% with only 9.1% of tokens.
Yifan Simon Liu, Ruifan Wu, Liam Gallagher et al.
Proposes Riemannian heat flow-based hypergraph neural network with adaptive local exchanger, addressing long-range dependencies in heterophilic and homophilic hypergraphs.
Li Sun, Ming Zhang, Wenxin Jin et al.
SWE-Hub unifies environment automation, scalable synthesis, and diverse task generation to support continuous, executable software engineering tasks.
Yucheng Zeng, Shupeng Li, Daxiang Dong et al.
AxProverBase achieves competitive performance with iterative proof refinement and context management in a simplified architecture.
Borja Requena, Austin Letson, Krystian Nowakowski et al.
PATRA employs pattern-aware alignment and balanced reinforcement learning to enhance time series question answering, achieving significant performance gains.
Junkai Lu, Peng Chen, Xingjian Wu et al.
SkillNet builds a unified skill ontology with multi-source creation and multi-dimensional evaluation, enhancing scalable AI skill management.
Yuan Liang, Ruobin Zhong, Haoming Xu et al.
Personalized LLM agents integrate user signals for cross-component interaction, enhancing long-term adaptability.
Yue Xu, Qian Chen, Zizhan Ma et al.
VeRO framework employs versioning and structured feedback to optimize coding agents, achieving up to 8% performance gains across tasks.
Varun Ursekar, Apaar Shanker, Veronica Chatrath et al.
NoRD achieves efficient vision-language-action learning on Waymo and NAVSIM with <60% data, no reasoning.
Ishaan Rawal, Shubh Gupta, Yihan Hu et al.
RB-VLA model excels in long-horizon tasks, achieving a 52.5% success rate improvement.
Vaidehi Bagaria, Bijo Sebastian, Nirav Kumar Patel
WarpRec unifies academic rigor and industrial scale with a backend-agnostic architecture, supporting 50+ algorithms and integrating CodeCarbon for energy tracking.
Marco Avolio, Potito Aghilar, Sabino Roccotelli et al.
Training AI agents in Corecraft high-fidelity environment, GLM 4.6 task pass rate improved from 25.37% to 36.76%.
Sushant Mehta, Logan Ritchie, Suhaas Garre et al.