GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.CV 2603.29281

PRISM: A Multi-View Multi-Capability Retail Video Dataset for Embodied Vision-Language Models

PRISM leverages multi-view videos and a 3D knowledge ontology to enhance embodied VLM's spatial, physical, and action understanding in retail environments.

Amirreza Rouhi, Parikshit Sakurikar, Satya Sai Reddy et al.

2026-03-31 35
cs.CV 2603.29194

Multi-Layered Memory Architectures for LLM Agents: An Experimental Evaluation of Long-Term Context Retention

Multi-Layer Memory Framework enhances LLMs' long-term context retention, achieving 46.85% success rate and 56.90% six-period retention.

Sunil Tiwari, Payal Fofadiya

2026-03-31 3
cs.CV 2603.29165

LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning

LatentPilot internalizes future visual dynamics via privileged supervision, achieving SOTA in VLN benchmarks with 66.3% success rate.

Haihong Hao, Lei Chen, Mingfei Han et al.

2026-03-31 46
cs.CV 2603.29163

SparseDriveV2: Scoring is All You Need for End-to-End Autonomous Driving

SparseDriveV2 leverages dense static vocabularies with trajectory decomposition, achieving 92.0 PDMS, challenging the necessity of dynamic proposals.

Wenchao Sun, Xuewu Lin, Keyu Chen et al.

2026-03-31 39
cs.CV 2603.29009

MEDiC: Multi-objective Exploration of Distillation from CLIP

MEDiC unifies CLIP distillation and pixel reconstruction, reaching 73.9% kNN and 85.1% fine-tuning accuracy on ImageNet-1K.

Konstantinos Georgiou, Maofeng Tang, Hairong Qi

2026-03-31 24
cs.CV 2603.28503

Bridging the Geometry Mismatch: Frequency-Aware Anisotropic Serialization for Thin-Structure SSMs

FGOS-Net achieves 91.3% mIoU and 97.1% clDice via frequency-geometric disentanglement.

Jin Bai, Huiyao Zhang, Qi Wen et al.

2026-03-30 40
cs.CV 2603.28088

GEMS: Agent-Native Multimodal Generation with Memory and Skills

GEMS framework enhances multimodal generation with memory and skills, enabling Z-Image-Turbo to surpass Nano Banana 2 on GenEval2.

Zefeng He, Siyuan Huang, Xiaoye Qu et al.

2026-03-30 40
cs.CV 2603.27931

A Cross-Scale Decoder with Token Refinement for Off-Road Semantic Segmentation

Proposed a Cross-Scale Decoder for off-road semantic segmentation, achieving 89.97 mIoU.

Seongkyu Choi Jhonghyun An

2026-03-30 44
cs.CV 2603.27773

RINO: Rotation-Invariant Non-Rigid Correspondences

RINO achieves rotation-invariant non-rigid matching via RINONet, significantly improving 3D shape correspondence accuracy.

Maolin Gao, Shao Jie Hu-Chen, Congyue Deng et al.

2026-03-30 0
cs.CV 2603.27206

Make It Up: Fake Images, Real Gains in Generalized Few-shot Semantic Segmentation

Syn4Seg framework enhances GFSS performance by generating diverse synthetic images and pseudo-labels, achieving significant improvements on PASCAL-5i and COCO-20i.

Guohuan Xie, Xin He, Dingying Fan et al.

2026-03-28 43
cs.CV 2603.27115

SJD-VP: Speculative Jacobi Decoding with Verification Prediction for Autoregressive Image Generation

SJD-VP enhances autoregressive image generation by predicting verification to improve acceleration and quality.

Bingqi Shan, Baoquan Zhang, Xiaochen Qi et al.

2026-03-28 38
cs.CV 2603.25931

DiReCT: Disentangled Regularization of Contrastive Trajectories for Physics-Refined Video Generation

DiReCT improves VideoPhy physical commonsense score by 16.7% without increasing training time.

Abolfazl Meyarian, Amin Karimi Monsefi, Rajiv Ramnath et al.

2026-03-27 6
cs.CV 2603.25827

Fus3D: Decoding Consolidated 3D Geometry from Feed-forward Geometry Transformer Latents

Fus3D leverages pretrained multi-view geometry transformer latents to directly regress dense SDF in under 3 seconds, enabling fast, pose-free 3D reconstruction.

Laura Fink, Linus Franke, George Kopanas et al.

2026-03-27 21
cs.CV 2603.25020

GDPO-Listener: Expressive Interactive Head Generation via Auto-Regressive Flow Matching and Group reward-Decoupled Policy Optimization

GDPO-Listener generates expressive head motions via Auto-Regressive Flow Matching and Group reward-Decoupled Policy Optimization.

Zhangyu Jin, Maksim Siniukov, Deuksin Kwon et al.

2026-03-26 4
cs.CV 2603.24793

AVControl: Efficient Framework for Training Audio-Visual Controls

AVControl leverages LTX-2 with LoRA adapters for efficient multi-modal audio-visual control, achieving high performance with minimal training steps.

Matan Ben-Yosef, Tavi Halperin, Naomi Ken Korem et al.

2026-03-26 35
cs.CV 2603.24581

Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving

Latent-WAM achieves efficient end-to-end autonomous driving with spatially-aware and dynamics-informed latent world representations, scoring 89.3 on NAVSIM v2.

Linbo Wang, Yupeng Zheng, Qiang Chen et al.

2026-03-26 4 citations 1068
cs.CV 2603.24577

EndoVGGT: GNN-Enhanced Depth Estimation for Surgical 3D Reconstruction

EndoVGGT enhances surgical 3D reconstruction with DeGAT, improving PSNR by 24.6% and SSIM by 9.1%.

Falong Fan, Yi Xie, Arnis Lektauers et al.

2026-03-26 239
cs.CV 2603.24575

VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models

VFIG uses vision-language models for complex figure-to-SVG conversion, achieving a VLM-Judge score of 0.829.

Qijia He, Xunmei Liu, Hammaad Memon et al.

2026-03-26 2 citations 212
cs.CV 2603.26790

Elucidating the Design Space of Flow Matching for Cellular Microscopy

Proposed a simplified, stable flow-matching generative model, scaled twice as large, with twofold FID reduction and improved unseen molecule simulation using molecular embeddings.

Charles Jones, Emmanuel Noutahi, Jason Hartford et al.

2026-03-25 23
cs.CV 2603.23501

MedObvious: Exposing the Medical Moravec's Paradox in VLMs via Clinical Triage

MedObvious exposes the Medical Moravec's Paradox in VLMs via a 1,880-task benchmark for clinical triage.

Ufaq Khan, Umair Nawaz, L D M S S Teja et al.

2026-03-25 2 citations 262
Prev 1 ... 34 35 36 37 38 39 40 ... 137 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home