Simple Supervision Is Hard to Beat: A Bitter Lesson from Sparse Target Labels in Domain-Adaptive Object Detection
RTSM method improves sparse-label domain-adaptive object detection by 1.7 to 18.3 AP50.
Lijun Zhang, Ruinian Xu, Mudit Agrawal
RTSM method improves sparse-label domain-adaptive object detection by 1.7 to 18.3 AP50.
Lijun Zhang, Ruinian Xu, Mudit Agrawal
HSAP integrates hierarchical sequence-aware parallelism with JIT-compiled SAP for ultra-long sequences, outperforming SOTA methods.
Songxin Zhang, Zejian Xie, Zhuoyang Song et al.
Arko-T is a 4B-parameter transformer that maps natural language directly into executable, editable parametric CAD programs, outperforming seven frontier LLMs on 12 metrics.
Liang Wang, Zhaoyang Xi, Zekai Xiang et al.
HUMEMBR combines continuous memory construction with structured retrieval, enabling long-term human routine modeling for improved robot reasoning.
Samira Huber, Klaas Pelzer, Duc M. Nguyen et al.
RenderFormer++ enhances global illumination rendering efficiency with physics-informed guidance and hierarchical tokenization.
Huangsheng Du, Haoran Zhu, Youcheng Cai et al.
FlowAWR performs advantage-weighted velocity rectification without SDEs or CFG, reaching PickScore 24.12 in 1.2k steps on SD3.5-Medium.
Zheming Fu, Ruizhe He, Wei Shang et al.
BrainJanus uses a unified autoregressive model with a neural tokenizer to enable bidirectional brain, vision, and language understanding and generation.
Haitao Wu, Qirui Zhang, Zhouheng Yao et al.
VisReflect enhances fine-grained perception in long visual contexts via latent visual reflection, achieving 4.1% improvement on image benchmarks and 1.8% on video benchmarks.
Xiaoqian Shen, Mohamed Elhoseiny
Develops KL divergence bounds for acceptance criteria in speculative decoding, applicable to greedy, relaxed, and tree-based decoding, enhancing practical inference reliability.
Aaryam Sharma
TACO optimizes tool calls in visual QA using DAPR and OGAR, enhancing accuracy.
Mingkuan Feng, Jinyang Wu, Hao Gu et al.
Clarus introduces a project-agent-resource model with a four-layer architecture for traceable, open scientific collaboration at web scale.
Zihan Guo, Zeyi Chen, Zhiyu Chen et al.
This study analyzes optimizer-dependent training dynamics via Hessian eigenvector displacement and localization, revealing SGD stabilizes directions while Adam induces eigenvector reorganization.
Marcelina Marjankowska, Valerio Modugno, Paolo Barucca
Introduced Blind Gap and Visual Gain metrics to reveal visual dependence issues in traffic accident VideoQA.
Sena Korkut, María Alejandra Bravo Sarmiento, Sanghwan Kim et al.
LatEnt Noise maSk (Lens) enhances multimodal large language models by reducing visual redundancy, improving VQA datasets by 2.4-6.4 points.
Kai Jiang, Ruishu Zhu, Siqi Huang et al.
Introduced STE and SPID frameworks to quantify semantic information flow and multi-source contributions in communication.
Leonardo S. Goodall, Andrea I. Luppi, Pedro A. M. Mediano
InnerZoom achieves efficient GUI grounding with a single forward pass, significantly improving benchmark scores.
Chen Liu, Ling Chen, Hanzhang Zhou et al.
VISTA enables self-managed context in LLMs via a training-free, model-agnostic layer that visualizes and archives internal states, boosting long-horizon task performance.
Binyan Xu, Haitao Li, Kehuan Zhang
Proposes posterior optimal E-values for non-convex parameter sets, applied to sequential voting system tests.
Adrienne Tuynman, Timothée Mathieu
FWS builds FaithfulQA from six VQA benchmarks; 60K faithful SFT raises RL accuracy and stabilizes visually grounded reasoning.
Peng, Lee, Yin Zhang et al.
IHDec uses JSD-based role attribution and contrastive decoding for real-time hierarchy correction, improving multi-turn instruction adherence by 11.98pp.
Nicole Geumheon Liu, Haeun Jang, Yonghyun Jun et al.