CORGI: Consistency-Aware 3D Dog Reconstruction from a Single Image in the Wild
CORGI employs CDOG, CA-3DGS, and DCGR to reconstruct high-fidelity 3D dogs from a single image without supervision.
Yuxiao Wu, Weile Li, Boyi Zhu et al.
CORGI employs CDOG, CA-3DGS, and DCGR to reconstruct high-fidelity 3D dogs from a single image without supervision.
Yuxiao Wu, Weile Li, Boyi Zhu et al.
PointSplat predicts Gaussian primitives directly from point clouds, enabling efficient real-time 3D human reconstruction with high quality.
Yujie Guo, Yudong Jin, Lingteng Qiu et al.
ERA introduces entropy-guided visual token pruning with bias rectification, effectively addressing attention logit collapse and boosting inference efficiency in multimodal large models.
Yuhao Wang, Mu Qiao, Haiwen Diao et al.
InstanceControl generates complex images without instance labeling, enhancing precision and control.
Xiaoyu Liu, Huan Wang, Fan Li et al.
MemLearner employs a learned query mechanism leveraging pre-trained visual priors to enhance scene memory in video world models, significantly improving scene consistency under occlusion and dynamic scenarios.
Jiwen Yu, Jianxiong Gao, Jianhong Bai et al.
Introduces WorldRoamBench, a benchmark for long-horizon stability of interactive world models, evaluating per-frame action accuracy, visual drift, physics, and memory over 600+ cases.
Ting-Bing Xu, Jiacheng Sui, Zhe Gao et al.
Localized Conformal Prediction (LCP) combined with vision-language models and nonlinear cosine similarity transformation reduces average prediction set size while maintaining coverage.
Clément Fuchs, Tim Bary, Benoît Macq
MV-GEL employs multi-view ranking and VLM reasoning to localize geometric entities on meshes, achieving up to 1.7× IoU improvement.
Kartik Bali, Roland Aydin
SimpleSearch-VL employs FAR for efficient sampling, integrates evidence verification for reliability, and maintains lightweight tools, significantly boosting multimodal agentic search performance.
Ming Dai, Zhihong Lu, Jinjie Gu et al.
Learning from Failure: Using failed trajectories for inference-time self-improvement, achieving a 6.6% success rate increase.
Xueqiao Sun, Xiaohan Wang, Ludwig Schmidt et al.
ForgeDrive achieves 90.3 EPDMS on NAVSIM by unifying driving simulation and planning via visual-action cross-conditioning.
Xuchang Zhong, He Zheng, Chenxu Zhao et al.
Proposes View-PNDF, a neuron-level fine-tuning method that enhances multi-view X-ray report consistency, achieving state-of-the-art results.
Yucheng Chen, Jinjing Zhu, Yang Yu et al.
Proposes Turing-inspired Label Imitation Game (LIG) with Transformer-based TTN for zero-shot pseudo-label pruning, improving detection F1 by up to 44%.
Brent A. Griffin, Jason J. Corso
RTSM method improves sparse-label domain-adaptive object detection by 1.7 to 18.3 AP50.
Lijun Zhang, Ruinian Xu, Mudit Agrawal
BrainJanus uses a unified autoregressive model with a neural tokenizer to enable bidirectional brain, vision, and language understanding and generation.
Haitao Wu, Qirui Zhang, Zhouheng Yao et al.
VisReflect enhances fine-grained perception in long visual contexts via latent visual reflection, achieving 4.1% improvement on image benchmarks and 1.8% on video benchmarks.
Xiaoqian Shen, Mohamed Elhoseiny
Introduced Blind Gap and Visual Gain metrics to reveal visual dependence issues in traffic accident VideoQA.
Sena Korkut, María Alejandra Bravo Sarmiento, Sanghwan Kim et al.
LatEnt Noise maSk (Lens) enhances multimodal large language models by reducing visual redundancy, improving VQA datasets by 2.4-6.4 points.
Kai Jiang, Ruishu Zhu, Siqi Huang et al.
InnerZoom achieves efficient GUI grounding with a single forward pass, significantly improving benchmark scores.
Chen Liu, Ling Chen, Hanzhang Zhou et al.
Nemotron-Labs-Diffusion-Image introduces a masked discrete diffusion model with token editing and GCE, achieving 0.90 on GenEval for high-res text-to-image synthesis.
Shufan Li, Greg Heinrich, Hanrong Ye et al.