StakeBench: Evaluating Language Understanding Grounded in Market Commitment
StakeBench evaluates language understanding via market commitment, linking 560,876 comments to four diagnostic tasks.
Yunhua Pei, Jingyu Hu, Yiwei Shi et al.
StakeBench evaluates language understanding via market commitment, linking 560,876 comments to four diagnostic tasks.
Yunhua Pei, Jingyu Hu, Yiwei Shi et al.
This paper introduces a three-level collapse model diagnosing why autonomous research systems lack true scientific closure, emphasizing external validation and community engagement.
Shuai Wang, Xinyuan Tian, Pangpang Liu et al.
Proposed CoAD fuses classification and reconstruction with probabilistic soft masking, boosting time series anomaly detection accuracy.
Qideng Tang, Dai Chaofan, Wubin Ma et al.
Introduces EQA-Decision dataset with over 4 million multimodal QA pairs, covering static scene, spatial understanding, task dynamics, and instant decision, for embodied reasoning.
Xicheng Gong, Qiwei Li, Peiran Xu et al.
FEP-Diff achieves cognitively plausible trajectory prediction under partial observability using the Free Energy Principle.
Yanping Wu, Ji Zhang, Hao Chen et al.
MIND model achieves 2.06 FID on ImageNet by explicitly modeling data manifold geometry.
Duoduo Xue, Zhiyu Zhu, Junhui Hou
DecoR optimizes LLM routing via query decomposition and historical matching, achieving higher accuracy and reduced inference costs.
Bo Lv, Jingbo Sun
TapSampling combines Action-VAE candidate sampling with task-progress verification, raising real-robot success from 78.3% to 83.3%.
Sizhe Zhao, Shengping Zhang, Shuo Yang et al.
BrickAnything uses geometry-conditioned autoregressive modeling with structure-aware tokenization to generate stable, geometrically faithful brick structures, reducing rollback by 70%.
Zhengyang Ni, Feng Yan, Yu Guo et al.
Point-via-text method resolves binding problem in vision language models, enhancing multi-object task performance.
Udith Haputhanthri, Declan Campbell, Rim Assouel et al.
Lightweight confidence-aware language model achieves SOTA success in autonomous driving decision-making, with low latency and high robustness.
Ruoyu Yao, Ruiguo Zhong, Pei Liu et al.
Proposes MixT, a Hamiltonian-inspired local operator approach, achieving over 50% parameter reduction in billion-parameter LLMs.
Ying Lu, Peng-Fei Zhou, Qi-Xuan Fang et al.
The study proposes an interference-aware experiment design framework, selecting designs in ads and recommendation systems, with a risk of 1.295 on Criteo ads.
Prashant Shekhar, Caroline Howard
Mimir is a 1.6B parameter multilingual concept model trained on 388 billion sentences, shifting from token to concept-level understanding.
Elio Musacchio, Lucia Siciliani, Pierpaolo Basile
Meta-Agent synthesizes and verifies task-specific multi-agent systems, reaching 82.7 average across six benchmarks.
Andy Xu, Yu-Wing Tai
Boosting inference with guided stochastic exploration: accuracy on Sudoku-Extreme improved from 85.9% to 98.0%.
Andrew Corbett, Archit Sood, Anna Tzatzopoulou et al.
ReWA combines reparameterization, weight decay, and adaptive learning rate to enhance sparse optimization, addressing instability and improving pruning in deep models.
Huangyu Xu, Jingqin Yang, Qianqian Xu et al.
This paper introduces robust diffusion operators and hidden-state re-propagation, significantly improving pre-propagation GNN performance on heterophilic graphs.
Zichao Yue, Zhiru Zhang
VEOcc uses a voxel-centric framework for online semantic occupancy prediction, excelling on Occ-ScanNet.
Ruoyu Wang, Yong Liu, Sheng Tao et al.
UTTSI scales inference by uncertainty, achieving 5.3% relative online CTR gain at about 2.8× average cost.
Moyu Zhang, Yun Chen, Yujun Jin et al.