ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes
ZipTok3D achieves high-fidelity 3D reconstruction with compact token prefixes, using only one token on ShapeNet.
Mingda Lin, Weijie Wang, Zeyu Zhang et al.
ZipTok3D achieves high-fidelity 3D reconstruction with compact token prefixes, using only one token on ShapeNet.
Mingda Lin, Weijie Wang, Zeyu Zhang et al.
This study mechanistically uncovers how LLMs evaluate summaries, revealing a two-stage process with attention-based error detection and MLP-based score crystallization at layers 25-26.
Himil Vasava, Ming Jiang
Proposes PTA-IRT, integrating trajectory data to improve software agent benchmarking with 10% calibration, outperforming traditional IRT.
Kefeng Duan, Dewu Zheng, Yanlin Wang et al.
ACToR identifies critical tokens during code generation and triggers on-demand retrieval, boosting accuracy by 8.4% on RepoExec and 15.4% on CoderEval.
Kefeng Duan, Dewu Zheng, Yanlin Wang et al.
CordisBench benchmarks LLMs' reasoning about component lifecycles in dynamic systems, with 1200 questions revealing performance drops as interactions increase.
Damien Sileo, Dimitri Kachler
Developed a mechanism framework using nested cyclical monotonicity for AI agents with unknown preferences and capabilities, enabling incentive-compatible control in multi-agent settings.
Dirk Bergemann, Andrew Koh, Stephen Morris
Introduces StudentSim framework, combining pooled pretraining and individual fine-tuning to enhance student simulators' fidelity and guidance responsiveness.
Ke Yang, Chenglong Wang, Michel Galley et al.
Using causal mixed-precision intervention, the study finds damage is widespread; global fine-grained quantization outperforms local layer repair by 21-52 points.
Jundong Hu, Shekar Ramachandran
SpatialGuard employs structured layout and verification to enhance spatial fidelity in complex 3D text-to-image generation, achieving state-of-the-art results.
Ziyun Qian, Zizhi Chen, Yizhou Liu et al.
SG-AMP combines robust depth completion, scene graph reasoning, and semantics-aware active view planning, achieving 55.27% semantic mIoU and 38.67% PQ on pepper data.
Rohit Menon, Shiva Rudra Lolla, Niklas Mueller-Goldingen et al.
A 35B MoE-based document understanding model with difficulty-aware data curation reduces deployment costs by over 80%.
Maksim Evdokimov, Matvey Ivanov, Dmitrii Tsiupin et al.
H3-World leverages MiniMax-H3 to enable precise, temporally grounded world control via natural language, with minimal fine-tuning and high generalization.
Danze Chen, Zeqing Wang, Ziyue Lin et al.
Introduces SCILAWS-BENCH, a real-data benchmark evaluating LLMs' ability to discover scientific laws, covering 118 problems, 291 candidate laws, and 8M data points.
Yiming Huang, Ziche Liu, Zhuohang Wu et al.
Layer-wise probing of V-JEPA 2 and VideoMAE-v2 reveals that camera motion is encoded mainly in mid-layers, forming smooth trajectories in feature space, with spline interpolation improving motion coherence.
Sharon S. Musa, Fereshteh Forghani, Harrish Thasarathan et al.
LRNBA architecture employs linear neural bases for parameter-efficient network compression, achieving comparable or superior performance with fewer parameters.
Binshuai Wang
NashDreamer employs centralized world models to efficiently converge to Nash equilibrium in two-player zero-sum imperfect-information games, improving early training sample efficiency by ~30%.
Tomáš Holeček, Viliam Lisý
Knowledge distillation during mid-training enhances reasoning but slows factual recall; proposes Switch Distillation.
Jacqueline He, Howard Yen, Shuyue Stella Li et al.
Gekko uses relative reconstruction error improvement to enhance 3D features without 3D labels, outperforming CroCo.
Thibaut Loiseau, Guillaume Bourmaud, Vincent Lepetit
VirSqueezer combines SenseGlove, MPM, and diffusion generation for fine-grained squeezing, but the paper excerpt reports no numeric benchmark values.
Qian Zhang, Xiaoming Chen, Xiaorui Ma et al.
A behavior architecture combining environment templates and behavior trees enables fast, resilient, and adaptable humanoid loco-manipulation with runtime editing.
Duncan Calvert, Luigi Penco, Dexton Anderson et al.