WALL-WM: Carving World Action Modeling at the Event Joints
WALL-WM enhances video-action learning with event-based Vision-Language-Action pretraining.
Shalfun Li, Victor Yao, Charles Yang et al.
WALL-WM enhances video-action learning with event-based Vision-Language-Action pretraining.
Shalfun Li, Victor Yao, Charles Yang et al.
MOSS-Video-Preview achieves real-time video understanding via cross-attention, boosting speed by 5x.
Pengyu Wang, Chenkun Tan, Shaojun Zhou et al.
DocFormFlow improves formatting accuracy by 72.53% while reducing token consumption.
Shihao Rao, Liang Li, Jiapeng Liu et al.
Proposes a unified, representation- and geometry-guided discrete tokenizer for autonomous driving, improving scene reconstruction and planning performance.
Ziyang Yao, Zeyu Zhu, YunCheng Jiang et al.
Set-Supervised Diffusion Policy learns action-chunking diffusion through human corrections, enhancing robotic manipulation performance.
Zhaoting Li, Gang Chen, Javier Alonso-Mora et al.
Proposes the 'Decan' metric, using single-pass log-probabilities to measure diversity in AI and human outputs, achieving 0.846 on McDiv benchmark.
Matthew Khoriaty, David Williams-King, Shi Feng
Trans2Occ predicts voxel occupancy of transparent objects from single-view RGB images, achieving 85% grasp success rate.
Yixuan Yang, Sha Zhang, Rui Li et al.
LLMs generate debate essays with only 3.4% unique main arguments, far below humans' 65.3%.
Yekyung Kim, Yapei Chang, Chau Minh Pham et al.
MINTS achieves optimal solutions for multi-armed bandit problems using a minimalist Bayesian framework, reaching the Lai-Robbins constant.
Kaizheng Wang
Introduced LongJudgeBench, a benchmark for evaluating LLMs as long-form judges, with an average output length over 9000 tokens, revealing significant instability.
Junjie Chen, Yuxi Dong, Haitao Li et al.
TLG improves video QA accuracy by 24.5% to 71.37% using temporal logic grounding.
Ali Alavi
Hierarchical object representation with points, meshes, superquadrics, and inflated superquadrics improves robot scene understanding and navigation accuracy.
Ceng Zhang, Wan Su, Mohamed Samshad et al.
Multi-Agent Computer Use (MACU) improves performance by 3.4-25.5% on desktop and web benchmarks.
Jing Yu Koh, Ruslan Salakhutdinov, Daniel Fried
This paper systematically evaluates 9 neuron models and 3 spike encoding schemes in SNN-based network intrusion detection, revealing latency encoding as the most effective approach, achieving 92.11% accuracy.
Raj Patel, David Amebley, Taye Akinrele et al.
Proposes structured post-retrieval assembly, boosting memory accuracy to 82%/93% single-hop and 27%/41% multi-hop, surpassing all prior results.
Vikas Reddy, Sumanth Reddy Challaram
Introduces a global boundary condition method (ZA) that reduces quantum resource costs in QLBM by avoiding segmentation, enabling complex geometry handling.
Călin A. Georgescu, Matthias Möller
BenchEvolver employs solution evolution to automatically generate harder, verifiable coding tasks, significantly reducing model success rates and enhancing discriminative evaluation.
Yangzhen Wu, Aaron J. Li, Wenjie Ma et al.
Andes introduces a self-evolving tree routing framework to enhance weak models' instruction alignment, achieving SOTA results with 33.39% accuracy on PostTrainBench.
Zhengyang Zhao, Shengjie Ye, Lu Ma et al.
The study proves that 'Domination-Avoiding' learning agents do not collude in markets, including mean-based and internal regret-minimizing algorithms.
Noam Nisan, Emmanuel Zerah
TrOPD employs trust-region strategies and multiple KL estimators to stabilize on-policy distillation, outperforming SOTA with +6.18 performance points.
Xingrun Xing, Haoqing Wang, Boyan Gao et al.