GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.CV 2506.08052

ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving

ReCogDrive combines VLM and diffusion planning to generate smooth, safe trajectories, achieving SOTA on NAVSIM with PDMS 90.8.

Yongkang Li, Kaixin Xiong, Xiangyu Guo et al.

2025-06-09 44
cs.CV 2506.06836

Harnessing Vision-Language Models for Time Series Anomaly Detection

Proposes VLM4TS with ViT4TS for zero-shot time series anomaly detection, achieving 24.6% F1 improvement and 36x token efficiency.

Zelin He, Sarah Alnegheimish, Matthew Reimherr

2025-06-07 13 citations 42
cs.CV 2506.05558

On-the-fly Reconstruction for Large-Scale Novel View Synthesis from Unposed Images

Proposes a real-time large-scale scene reconstruction method combining fast initial pose estimation and direct Gaussian primitive sampling, completing scene modeling within 30 minutes with PSNR 21.7dB.

Andreas Meuleman, Ishaan Shah, Alexandre Lanvin et al.

2025-06-06 43 citations 33
cs.CV 2506.05312

Do It Yourself: Learning Semantic Correspondence from Pseudo-Labels

Proposes 3D-aware pseudo-label learning for semantic correspondence, achieving over 4% improvement on SPair-71k.

Olaf Dünkel, Thomas Wimmer, Christian Theobalt et al.

2025-06-06 43
cs.CV 2506.05302

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

PAM combines SAM 2 and LLMs for efficient, multi-task regional understanding in images and videos, achieving 1.2-2.4× faster performance.

Weifeng Lin, Xinyu Wei, Ruichuan An et al.

2025-06-06 32
cs.CV 2506.04590

Follow-Your-Creation: Empowering 4D Creation through Video Inpainting

Follow-Your-Creation employs a video inpainting framework to generate and edit 4D content from monocular videos, achieving multi-view consistency and flexible editing.

Yue Ma, Kunyu Feng, Xinhua Zhang et al.

2025-06-05 23
cs.CV 2506.04220

Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs

Struct2D enables 3D spatial reasoning using perception-guided structured 2D inputs, achieving high performance without explicit 3D data.

Fangrui Zhu, Hanhui Wang, Yiming Xie et al.

2025-06-05 44
cs.CV 2506.04122

Contour Errors: Ego-Centric Matching for 3D Multi-Object Tracking Performance Evaluation

Proposes Contour Errors (CE), an ego-centric matching metric based on Hausdorff distance, improving 3D MOT evaluation over IoU and CPD.

Sharang Kaul, Simon Bultmann, Mario Berk et al.

2025-06-05 42
cs.CV 2506.04034

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning

Rex-Thinker uses Chain-of-Thought reasoning for object referring, enhancing precision and interpretability.

Qing Jiang, Xingyu Chen, Zhaoyang Zeng et al.

2025-06-04 31
cs.CV 2506.01908

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Temporal-RLT framework enhances video understanding efficiency through reward design and data selection, significantly reducing training data.

Hongyu Li, Songhao Han, Yue Liao et al.

2025-06-03 14
cs.CV 2506.01853

ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding

ShapeLLM-Omni employs 3D VQVAE and autoregressive transformers to enable unified 3D understanding and generation, supported by the 3D-Alpaca dataset.

Junliang Ye, Zhengyi Wang, Ruowen Zhao et al.

2025-06-03 39
cs.CV 2506.01802

UMA: Ultra-detailed Human Avatars via Multi-level Surface Alignment

UMA introduces multi-level surface alignment guided by 2D point tracking, achieving 30% reduction in geometric error and 20% detail enhancement at 4K resolution.

Heming Zhu, Guoxing Sun, Christian Theobalt et al.

2025-06-02 31
cs.CV 2506.00742

ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary

ArtiScene leverages image intermediary, enabling training-free, text-driven 3D scene generation with high aesthetic quality.

Zeqi Gu, Yin Cui, Zhaoshuo Li et al.

2025-06-01 58
cs.CV 2505.24718

Reinforcing Video Reasoning with Focused Thinking

TW-GRPO framework enhances video reasoning with focused thinking, achieving 50.4% accuracy on CLEVRER.

Jisheng Dang, Jingze Wu, Teng Wang et al.

2025-05-30 14
cs.CV 2505.24705

RT-X Net: RGB-Thermal cross attention network for Low-Light Image Enhancement

Proposes RT-X Net, a cross-attention transformer that fuses RGB and thermal images, achieving superior low-light enhancement with PSNR of 27.75dB.

Raman Jha, Adithya Lenka, Mani Ramanagopal et al.

2025-05-30 33
cs.CV 2505.24625

Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors

Proposes VG LLM, leveraging a pre-trained 3D geometry encoder to learn spatial priors from videos, significantly improving 3D scene understanding and reasoning.

Duo Zheng, Shijia Huang, Yanyang Li et al.

2025-05-30 40
cs.CV 2505.23922

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding

ScaleLong evaluates long video understanding across timescales, revealing a U-shaped performance curve.

David Ma, Huaqing Yuan, Xingjian Wang et al.

2025-05-30 13
cs.CV 2505.23287

GenCAD-Self-Repairing: Feasibility Enhancement for 3D CAD Generation

Proposed GenCAD-Self-Repairing enhances CAD generation feasibility from 10% to 97% using diffusion guidance and self-repair, fixing 65% of infeasible designs.

Chikaha Tsuji, Enrique Flores Medina, Harshit Gupta et al.

2025-05-29 59
cs.CV 2505.23158

LODGE: Level-of-Detail Large-Scale Gaussian Splatting with Efficient Rendering

LODGE introduces a hierarchical LOD framework for large-scale 3D Gaussian Splatting, reducing GPU memory by 50% and increasing rendering speed to 200 FPS.

Jonas Kulhanek, Marie-Julie Rakotosaona, Fabian Manhardt et al.

2025-05-29 28
cs.CV 2505.22654

VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models

VScan accelerates LVLM inference by 2.91× with only 4.6% performance loss, using a two-stage token pruning strategy.

Ce Zhang, Kaixin Ma, Tianqing Fang et al.

2025-05-29 27
Prev 1 ... 61 62 63 64 65 66 67 ... 138 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home