Youtu-Parsing: Perception, Structuring and Recognition via High-Parallelism Decoding
Youtu-Parsing achieves 5-11x speedup via high-parallelism decoding, setting SOTA performance on OmniDocBench and olmOCR-bench.
Haoyu Cao, Kun Yin, Yunfei Wu et al.
Youtu-Parsing achieves 5-11x speedup via high-parallelism decoding, setting SOTA performance on OmniDocBench and olmOCR-bench.
Haoyu Cao, Kun Yin, Yunfei Wu et al.
This study empirically shows that AGENTS.md files reduce AI coding agents' token usage by 20% and task completion time by over 20% on GitHub pull requests.
Jai Lal Lulla, Seyedmoein Mohsenimofidi, Matthias Galster et al.
HINT introduces a hierarchical interaction autoregressive diffusion framework, achieving a low FID of 3.100 and high semantic coherence for multi-human motion generation.
Mengge Liu, Yan Di, Gu Wang et al.
OmegaUse achieves 96.3% on ScreenSpot-V2 using a Mixture-of-Experts model.
Le Zhang, Yixiong Xiao, Xinjiang Lu et al.
AMA framework achieves adaptive memory via multi-agent collaboration, reducing token consumption by 80%.
Weiquan Huang, Zixuan Wang, Hehai Lin et al.
Study shows effective dimension of representation geometry strongly predicts deep neural network generalization.
Sumit Yadav
Proposed TRACE benchmark with contrastive analysis detects reward hacking in code environments; GPT-5.2 achieves 63% detection rate.
Darshan Deshpande, Anand Kannappan, Rebecca Qian
Introduces Self-Distillation Fine-Tuning (SDFT), enabling continual learning with reduced catastrophic forgetting and improved new-task accuracy.
Idan Shenfeld, Mehul Damani, Jonas Hübotter et al.
Keel integrates Highway connections into Post-LayerNorm, enabling stable training of Transformer models over 1000 layers, outperforming Pre-LN in stability and expressivity.
Chen Chen, Lai Wei
VGGT-SLAM 2.0 employs a novel factor graph and attention-based verification to achieve real-time dense scene reconstruction, reducing pose error by 23%.
Dominic Maggio, Luca Carlone
ULEE combines unsupervised pretraining, self-generated goals, and meta-learning to enhance exploration and adaptation, outperforming baselines.
Octavio Pappalardo
Using deep autoencoders (DAE) to infer the intrinsic dimensionality of FPUT trajectories, finding ID≈2 in weak nonlinearity and ID=3 at β=1.1.
Gionni Marchetti
MDGR uses masked diffusion for SID generation, improving over SOTA by up to 10.78% and increasing online revenue by 1.20%.
Lingyu Mu, Hao Deng, Haibo Xing et al.
PROTEUS uses Lagrangian RL for SLA-aware multi-LLM routing, achieving 94.0% accuracy and 89.8% cost savings.
Amit Singh Bhatti, Vishal Vaddina, Dagnachew Birru
DeFM employs self-supervised pretraining on 60M depth images, learning geometric and semantic features for robotic tasks with state-of-the-art results.
Manthan Patel, Jonas Frey, Mayank Mittal et al.
Proposes OPSD, a self-distillation method where a single large model acts as both teacher and student, using known correct answers to improve reasoning efficiency.
Siyan Zhao, Zhihui Xie, Mengchen Liu et al.
K-Myriad maximizes collective state entropy across multiple policies, enhancing exploration in high-dimensional continuous RL tasks.
Vincenzo De Paola, Mirco Mutti, Riccardo Zamboni et al.
GenAgent employs agentic multimodal reasoning with tool invocation, achieving +23.6% performance on GenEval++.
Kaixun Jiang, Yuzheng Wang, Junjie Zhou et al.
Analyzed 14.8 million prompts to assess how different demographic cues affect LLM responses, revealing inconsistent effects across cues.
Manuel Tonneau, Neil K. R. Seghal, Niyati Malhotra et al.
Proposes StreamLoD-GS, a hierarchical LoD-based 3D Gaussian Splatting framework for real-time sparse-view video reconstruction, achieving PSNR 22.73dB with 0.2MB storage.
Xinhui Liu, Can Wang, Lei Liu et al.