Training, Reading, and Editing Legible Transformers
Proposes training legible transformers with variance floor and sparse units, achieving 78% detection and 50% attention channel legibility.
Mark Oskin
Proposes training legible transformers with variance floor and sparse units, achieving 78% detection and 50% attention channel legibility.
Mark Oskin
AgenticFocus converts human FPV videos into robot-trainable demonstrations with a SPARC score of -5.18.
Iaroslav Kolomiets, Miguel Altamirano Cabrera, Artem Lykov et al.
Proposes S2AE with spatial and semantic regularization, improving cross-modal concept alignment by 6.06%.
Weiduo Liao, Yunqiao Yang, Ying Wei
Procrustes-conditioned joint end-to-end Top-K sparse autoencoder extracts universal features across independent BERT seeds, achieving Pearson r≥0.70.
Bendegúz Váradi, Zoltán Kmetty
Introduces self-patching to diagnose and recover 58-75% of reasoning failure caused by knowledge-circuit misalignment in fine-tuned LLMs.
Lu Dai, Ziyang Rao, Yili Wang et al.
WCog-VLA integrates semantic forecasting and generative modeling, achieving 92.9 PDMS on NAVSIM, enabling proactive autonomous driving.
Xuerun Yan, Zhexi Lian, Nuoheng Zhang et al.
MobiDiff uses multi-channel discrete diffusion to efficiently generate privacy-preserving human mobility data.
Rongchao Xu, Lin Jiang, Dahai Yu et al.
AnyDexRT enables calibration-free high-quality dexterous hand retargeting with few-shot human guidance.
Chenxi Wang, Ying Feng, Hongjie Fang et al.
Introduces Temporal Ratio (TR) to quantify and improve video action model generalization, boosting success rates from 71.7% to 83.3% in robotic tasks.
Utkarsh A. Mishra, Yongxin Chen, Danfei Xu et al.
KronQ employs Kronecker-factored Hessian approximation with gradient covariance, achieving state-of-the-art 2-bit quantization, e.g., perplexity 7.93 on LLaMA-3-70B.
Donghyun Lee, Yuhang Li, Ruokai Yin et al.
DeepSearch-Evolve achieves self-distillation in DeepSearch-World, reaching 93.4% on HotpotQA.
Xinyu Geng, Xuanhua He, Sixiang Chen et al.
DreamCharacter-1 enhances 3D character generation with geometry and texture post-training, surpassing existing methods.
Weizhe Liu, Yunjie Wu, Xiangqian Shu et al.
Study finds LLM-generated skills do not significantly improve performance in data science workflows.
Wei-Jung Huang
DeLS-Spec improves inference speed and acceptance length by combining long and short contexts.
Hong-Kai Zheng, Piji Li
EvoSOP enables self-evolving LLM agents by iteratively synthesizing atomic actions into reusable SOPs, boosting success rates by 3-13% and reducing interaction rounds.
Haipeng Ding, Yuexiang Xie, Zhewei Wei et al.
Proposes a circuit-based mechanistic interpretability framework using SAEs and transcoders to disentangle polysemantic features in Transformer models.
Pranav Sawant, Jakub Krejčí
Introduces Projected Energy Matching using Helmholtz decomposition, reducing rotational artifacts in 3D energy models for medical image reconstruction.
Daniel Barco, Michal Balcerak, Suprosanna Shit et al.
PUF framework improves relationship recall by 18.1% on 3DSSG while maintaining real-time performance.
Yi Yang, Myrna Castillo, Bodo Rosenhahn et al.
Self-supervised pretraining with Point-M2AE enhances cross-site and cross-scale robustness in point cloud leaf-wood segmentation, boosting IoU by ~10% and improving downstream volume estimation.
Heeju Mun, Tackang Yang, Yunsoo Nam et al.
Activation-space anomaly detection in MAS outperforms graph-based methods, achieving F1 of 0.94/0.93 in synchronous/asynchronous settings.
Haowen Xu, Xue Tan, Lei Ma et al.