GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.CV 2506.15442

Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material

Hunyuan3D 2.1 generates high-fidelity 3D assets using Hunyuan3D-DiT and Hunyuan3D-Paint, enhancing geometric detail and material quality.

Team Hunyuan3D, Shuhui Yang, Mingxin Yang et al.

2025-06-18 32
cs.CV 2506.14907

PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning

Proposes PeRL, a permutation-enhanced RL framework, achieving state-of-the-art on multi-image reasoning benchmarks with significant margin.

Yizhen Zhang, Yang Ding, Shuoshuo Zhang et al.

2025-06-18 25
cs.CV 2506.14404

Causally Steered Diffusion for Automated Video Counterfactual Generation

Proposes CSVC, a causal prompt optimization framework guiding diffusion models to generate causally consistent counterfactual videos, improving effectiveness and quality.

Nikos Spyrou, Athanasios Vlontzos, Paraskevas Pegios et al.

2025-06-17 39
cs.CV 2506.14096

Image Segmentation with Large Language Models: A Survey with Perspectives for Intelligent Transportation Systems

Integrating LLMs like GPT-3 with vision models enhances traffic scene segmentation, achieving 85% mIoU and 78% zero-shot accuracy on Cityscapes and BDD100K.

Sanjeda Akter, Ibne Farabi Shihab, Anuj Sharma

2025-06-17 39
cs.CV 2506.13697

Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry

Vid-CamEdit enables video camera trajectory editing via generative rendering and geometry estimation, enhancing novel view video synthesis quality.

Junyoung Seo, Jisang Han, Jaewoo Jung et al.

2025-06-17 27
cs.CV 2506.13387

TR2M: Transferring Monocular Relative Depth to Metric Depth with Language Descriptions and Dual-Level Scale-Oriented Contrast

TR2M leverages image and text to predict pixel-wise scale maps, converting relative to absolute depth with high cross-domain generalization.

Beilei Cui, Yiming Huang, Long Bai et al.

2025-06-16 70
cs.CV 2506.12251

Efficient Multi-Camera Tokenization with Triplanes for End-to-End Driving

Proposes a triplane-based multi-camera tokenization, reducing tokens by 72% and doubling inference speed, for end-to-end autonomous driving.

Boris Ivanovic, Cristiano Saltori, Yurong You et al.

2025-06-14 41
cs.CV 2506.14827

DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning

DAVID-XR1 integrates fine-grained defect annotations and chain-of-thought reasoning, enabling interpretable AI-generated video detection with strong cross-generator generalization.

Yifeng Gao, Yifan Ding, Hongyu Su et al.

2025-06-13 8 citations 44
cs.CV 2506.11661

Prohibited Items Segmentation via Occlusion-aware Bilayer Modeling

Proposed occlusion-aware bilayer mask decoder with SAM integration achieves 94.4% mAP on prohibited item segmentation in X-ray images.

Yunhan Ren, Ruihuang Li, Lingbo Liu et al.

2025-06-13 24
cs.CV 2506.10977

QuadricFormer: Scene as Superquadrics for 3D Semantic Occupancy Prediction

QuadricFormer leverages superquadrics for efficient 3D occupancy prediction, achieving 21.11 mIoU with fewer primitives and lower computational cost.

Sicheng Zuo, Wenzhao Zheng, Xiaoyong Han et al.

2025-06-13 37
cs.CV 2506.10941

VINCIE: Unlocking In-context Image Editing from Video

VINCIE model learns in-context image editing from videos, achieving state-of-the-art results on multi-turn editing benchmarks.

Leigang Qu, Feng Cheng, Ziyan Yang et al.

2025-06-13 8
cs.CV 2506.10890

CreatiPoster: Towards Editable and Controllable Multi-Layer Graphic Design Generation

CreatiPoster employs a multi-layer protocol model to generate editable graphic compositions, with background synthesis ensuring visual harmony and flexibility.

Dexiang Hong, Zhao Zhang, Weidong Chen et al.

2025-06-13 41
cs.CV 2506.10857

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos

VRBench evaluates multi-step reasoning in long videos, with 960 videos and 8,243 QA pairs.

Jiashuo Yu, Yue Wu, Meng Chu et al.

2025-06-13 13
cs.CV 2506.10821

VideoExplorer: Think With Videos For Agentic Long-Video Understanding

VideoExplorer employs dynamic reasoning and temporal grounding to outperform baselines in long-video understanding, achieving 54.3% accuracy.

Huaying Yuan, Zheng Liu, Junjie Zhou et al.

2025-06-12 38
cs.CV 2506.09989

Hearing Hands: Generating Sounds from Physical Interactions in 3D Scenes

Rectified flow turns 3D hand trajectories into interaction sounds that are often indistinguishable from real audio.

Yiming Dou, Wonseok Oh, Yuqing Luo et al.

2025-06-12 32
cs.CV 2506.09980

Efficient Part-level 3D Object Generation via Dual Volume Packing

Proposes dual volume packing for efficient part-level 3D generation from a single image, avoiding segmentation priors.

Jiaxiang Tang, Ruijie Lu, Zhaoshuo Li et al.

2025-06-12 37
cs.CV 2506.08015

4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular Videos

Proposes 4D Gaussian Transformer (4DGT) for real-time dynamic scene reconstruction from monocular videos, significantly improving speed and accuracy.

Zhen Xu, Zhengqin Li, Zhao Dong et al.

2025-06-10 43 citations 35
cs.CV 2506.08009

Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion

Self Forcing introduces autoregressive training with distribution matching, enabling real-time video generation with high quality and low latency, outperforming traditional slow diffusion models.

Xun Huang, Zhengqi Li, Guande He et al.

2025-06-10 533 citations 36
cs.CV 2506.08005

ZeroVO: Visual Odometry with Minimal Assumptions

ZeroVO achieves zero-shot generalization across environments, improving performance by over 30%.

Lei Lai, Zekai Yin, Eshed Ohn-Bar

2025-06-10 16
cs.CV 2506.07491

SpatialLM: Training Large Language Models for Structured Indoor Modeling

SpatialLM integrates point cloud encoding with large language models, achieving state-of-the-art indoor scene layout estimation and 3D detection.

Yongsen Mao, Junhao Zhong, Chuan Fang et al.

2025-06-09 33
Prev 1 ... 60 61 62 63 64 65 66 ... 138 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home