Scaling Adaptive Depth with Norm-Agnostic Residual Networks
NAG architecture separates magnitude and direction in residual streams to address norm growth in deep models.
Tomás Figliolia, Beren Millidge
NAG architecture separates magnitude and direction in residual streams to address norm growth in deep models.
Tomás Figliolia, Beren Millidge
Proposes GhostPrint, a parameter-efficient attack framework that enables weak models to mimic strong LLMs, exposing fingerprint spoofing vulnerabilities.
Jiahao Zhang, Xiuyu Li, Suhang Wang
VinQA introduces two visual encoding methods, significantly improving visual citation accuracy in long-form multimodal document QA.
Young Rok Jang, Hyesoo Kong, Kyunghwan An et al.
Introduces active learning with low-rank structure for data selection, enhancing training efficiency.
Vincent Cohen-Addad, Sasidhar Kunapuli, Vahab Mirrokni et al.
SAG employs SQL-driven dynamic hyperedges, boosting multi-hop reasoning with 80% Recall@5 on MuSiQue, surpassing prior methods.
Yuchao Wu, Junqin Li, XingCheng Liang et al.
This paper provides the first runtime analysis of Cartesian Genetic Programming (CGP) in evolving Boolean functions, establishing bounds of O(nD^5) for conjunctions and exponential time for XOR, highlighting the impact of selection strategies.
Duc-Cuong Dang, Roman Kalkreuth, Andre Opris
Deep-VRM achieves full-spectrum forensic signal perception via residual injection, reaching 97.42% accuracy on GenImage dataset.
Kaiqing Lin, Zhiyuan Yan, Ruoxin Chen et al.
Proposes amortized neural network mapping for mean-shift particles, reducing Bayesian integral error in inverse problems.
Ali Siahkoohi
Metis employs decoupled video generation and action prediction with Mixture-of-Transformers, achieving state-of-the-art autonomous driving performance.
Jingyu Li, Zhe Liu, Dongnan Hu et al.
AIChilles finds 49 hidden regressions across 30 AI-evolved system programs using differential, constraint-aware search.
Yajie Zhou, Ao Li, Ashwin Silla et al.
LaWAM uses latent visual subgoals for efficient scene prediction, achieving 98.6% success with 187ms inference.
Jialei Chen, Kai Wang, Kang Chen et al.
Retrieval-augmented policy enables zero-shot extension of vision-language-action models at test time, reducing costs by replacing fine-tuning with retrieval.
Jeongeun Park, Juhan Park, Taekyung Kim et al.
Using SAM3 with structured prompts enables zero-shot spacecraft component segmentation without model fine-tuning, achieving 0.385 [email protected] on unseen satellites.
Nicholas A. Welsh, Lennon J. Shikhman, Monty Nehru Attazs et al.
Proposes TAG, converting egocentric videos into structured temporal graphs for zero-shot action recognition using VLMs, outperforming traditional pixel-based methods.
Bessie Dominguez-Dager, Francisco Gomez-Donoso, Miguel Cazorla et al.
Proposes a non-uniform timestep rescheduling method for diffusion inversion, reducing errors and improving image reconstruction accuracy by leveraging error analysis and dynamic programming.
Shangquan Sun, Ting Gong, Zhirui Liu et al.
DiRecT employs receding-horizon denoising in diffusion models, enforcing constraints only at the final trajectory, significantly improving safety and performance.
Paolo Giaretta, Zeyang Li, Navid Azizan
CausalDrive employs a real-time causal autoregressive world model with flow-matching and self-distillation, achieving 12 FPS interactive autonomous driving simulation without future layout conditioning.
Tianyi Yan, Huan Zheng, Dubing Chen et al.
Proposes IG-DOE, a large language model-assisted cooperative operator ensemble evolution algorithm, integrating multi-operator switching to significantly improve permutation flow shop scheduling performance.
Rui Xu, Yufan Liao, Haoze Lv et al.
CoMET-Agent achieves conditional multi-event temporal grounding in long videos, improving [email protected] by 6.1%.
Yuanhao Zou, Arthad Kulkarni, Lucas Tonanez et al.
TraCS integrates neuro-symbolic reasoning with motion prediction, improving accuracy and interpretability in multi-modal ground mobility.
Simon Kohaut, Felix Divo, Julius Hahnewald et al.