Complexity as Advantage: A Regret-Based Perspective on Emergent Structure
Introduces Complexity as Advantage (CAA) framework, defining complexity via predictive regret.
Oshri Naparstek
Introduces Complexity as Advantage (CAA) framework, defining complexity via predictive regret.
Oshri Naparstek
Introduces soft-constrained optimal transport to approximate Knothe-Rosenblatt maps, with proven convergence guarantees.
Ricardo Baptista, Franca Hoffmann, Minh Van Hoang Nguyen et al.
Proposes 'Thinking with Video' paradigm using Sora-2 to unify multimodal reasoning, achieving 92% accuracy on MATH and 69.2% on MMMU benchmarks.
Jingqi Tong, Yurong Mou, Hangcheng Li et al.
Batch prompting effectively suppresses overthinking in large reasoning models, reducing reasoning tokens by 76%.
Saurabh Srivastava, Janit Bidhan, Hao Yan et al.
This paper compares Leinster-Cobbold-Reeve (LCR) and Vendi Score (VS) for similarity-sensitive entropy across 53 large datasets.
Phuc Nguyen, Josiah Couch, Rahul Bansal et al.
LiveTradeBench evaluates LLMs in real-time multi-asset trading, revealing gaps between static benchmarks and actual performance.
Haofei Yu, Fenghai Li, Jiaxuan You
MUTANT employs a two-stage BPE training with language-aware preprocessing, reducing fertility by 39.5%, boosting inference throughput by 44%.
Souvik Rana, Arul Menezes, Ashish Kulkarni et al.
miniF2F-v2 significantly improves theorem proving accuracy to 70%.
Azim Ospanov, Farzan Farnia, Roozbeh Yousefzadeh
PublicAgent employs a multi-agent architecture based on LLMs for open data analysis, validating five core design principles.
Sina Montazeri, Yunhe Feng, Kewei Sha
TWIST2 integrates PICO4U VR and a custom 2-DoF neck to enable scalable, mocap-free humanoid teleoperation and data collection, achieving near 100% success in 100 demonstrations within 15 minutes.
Yanjie Ze, Siheng Zhao, Weizhuo Wang et al.
Lookahead Unmasking (LookUM) improves diffusion language decoding by path search, reducing errors with only 2-3 candidate paths, outperforming greedy methods.
Sanghyun Lee, Seungryong Kim, Jongho Park et al.
DoFlow employs continuous normalizing flows (CNF) integrated with causal DAGs for observational, interventional, and counterfactual time-series forecasting.
Dongze Wu, Feng Qiu, Yao Xie
Proposes a structured world model framework based on Hidden Markov Models (HMM) and switching linear dynamical systems (sLDS), supporting passive prediction and active control.
Lancelot Da Costa, Sanjeev Namjoshi, Mohammed Abbas Ansari et al.
Wonder3D++ employs cross-domain diffusion to generate high-fidelity 3D meshes from a single image, integrating multi-view normal maps and color images with attention mechanisms.
Yuxiao Yang, Xiao-Xiao Long, Zhiyang Dou et al.
AlphaEvolve combines LLM and automated evaluation in evolutionary search, discovering novel mathematical constructions across 67 problems, often surpassing known solutions.
Bogdan Georgiev, Javier Gómez-Serrano, Terence Tao et al.
Proposes Viewpoint Learning with a two-stage fine-tuning strategy and Viewpoint-100K dataset to activate spatial reasoning in multimodal large models, achieving significant improvements.
Xiaoyu Zhan, Wenxuan Huang, Hao Sun et al.
ExplicitLM introduces a million-scale external memory bank with a two-stage differentiable retrieval, achieving 43.67% improvement on knowledge tasks.
Chengzhang Yu, Zening Lu, Chenyang Zheng et al.
LiCoMemory uses hierarchical CogniGraph for lightweight, structured long-term memory, boosting reasoning efficiency by 23%.
Zhengjun Huang, Zhoujin Tian, Qintian Guo et al.
Using GPT-4o model enhances audio comprehension and generation, significantly improving audio interaction performance.
Siyin Wang, Zengrui Jin, Changli Tang et al.
MotionStream enables real-time, 29FPS streaming video generation on a single GPU, integrating motion control with long-sequence stability via causal distillation.
Joonghyuk Shin, Zhengqi Li, Richard Zhang et al.