Task-Agnostic Noisy Label Detection via Standardized Loss Aggregation
Standardized Loss Aggregation (SLA) detects noisy labels by aggregating standardized fold validation losses, outperforming hard-count baselines.
Inhyuk Park, Doohyun Park
Standardized Loss Aggregation (SLA) detects noisy labels by aggregating standardized fold validation losses, outperforming hard-count baselines.
Inhyuk Park, Doohyun Park
DriveFuture explicitly conditions current planning on predicted future latent states, achieving 55.5 EPDMS on NAVSIM-v2 navhard, surpassing previous SOTA.
Yufeng Hong, Xiaotian Zhou, Yingyan Li et al.
Re-Prefill introduces attention-guided second prefill to improve GUI element localization, boosting accuracy up to 4.3%.
Jiaping Lin, Fei Shen, Junzhe Li et al.
Proposes 'Imagining in 360°' framework with probabilistic spatial priors, boosting humanoid visual search success by over 30%.
Jingdong Zhang, Yizhou Wang, Zhengzhong Tu et al.
HyRo uses hyperbolic rotation to decouple hierarchical and semantic alignment, achieving state-of-the-art open-vocabulary semantic segmentation.
Hoang M. Truong, Hai Nguyen-Truong, Dang Huynh
ZAYA1-VL-8B integrates vision-specific LoRA adapters and bidirectional attention, boosting multimodal performance efficiently.
Hassan Shapourian, Kasra Hejazi, Olabode M. Sule et al.
STARFlow2 integrates language models and normalizing flows for unified multimodal generation, showing strong performance.
Ying Shen, Tianrong Chen, Yuan Gao et al.
Enhanced monocular depth prediction using distance transform over pre-semantic contours, significantly improving performance in low-texture areas.
Marwane Hariat, Antoine Manzanera, David Filliat
ForgeVLA learns VLA models from distributed vision-action pairs using federated learning, without language annotations, significantly improving performance.
Yuhao Zhou, Yunpeng Zhu, Yang Zhou et al.
Proposed Residual Latent Action (RLA) method using DINO residuals to improve world model prediction accuracy and reduce computational costs in multi-task robotics.
Xinyu Zhang, Zhengtong Xu, Yutian Tao et al.
SAFE challenge employs deep learning with pre-trained vision backbones and autoencoding to detect synthetic videos, achieving an average AUC of 0.82 across 13 models.
Kirill Trapeznikov, Gabriel Mancino-Ball, Jonathan Li et al.
HumanNet dataset enables Qwen model to substitute human video for robot data, enhancing learning efficiency.
Yufan Deng, Daquan Zhou
Prologue introduces prologue tokens to decouple generation and reconstruction, reducing gFID from 21.01 to 10.75 on ImageNet 256×256.
Bowen Zheng, Weijian Luo, Guang Yang et al.
Proposes InfoCoordiBridge, a neuro-symbolic framework integrating multi-sensor coordination and verifiable reasoning, boosting autonomous scene understanding reliability.
Shuo Liu, Lei Shi, Haowen Liu et al.
FSTM employs a two-stage training process with a single SDF to achieve indoor scene geometry and semantic reconstruction, improving speed by 2.3×.
Remi Chierchia, Léo Lebrat, David Ahmedt-Aristizabal et al.
Proposes CAFe-DINO, leveraging DINOv3 for open-vocabulary remote sensing segmentation without fine-tuning, outperforming supervised models.
Ryan Faulkenberry, Saurabh Prasad
Hyp2Former employs hyperbolic hierarchical embeddings to improve open-set panoptic segmentation, achieving state-of-the-art results.
Yao Lu, Rohit Mohan, Florian Drews et al.
NTIRE 2026 Challenge introduces RetinexFormerRefine, achieving SSIM of 0.5654 with models under 1MB for low-light image enhancement.
Jiebin Yan, Chenyu Tu, Weixia Zhang et al.
MeshReGen framework regenerates 3D objects from 2D images and initial 3D shapes using VecSet mechanism, enhancing detail consistency.
Geon Yeong Park, Roman Shapovalov, Rakesh Ranjan et al.
Proposes DynamicGUIBench and DynamicUI, leveraging video-based perception to address partial observability in high-dynamic GUI environments, achieving 20% performance gains.
Enqi Liu, Liyuan Pan, Zhi Gao et al.