GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.CV 2605.13366

Neural Surrogate Forward Modelling For Electrocardiology Without Explicit Intracellular Conductivity Tensor

Deep learning model predicts ECG with R2 of 0.949, reducing structural uncertainty.

Shaheim Ogbomo-Harmitt, Cesare Magnetti, Jakub Grzelak et al.

2026-05-13 0
cs.CV 2605.16405

Concepts Worth Having: Refining VLM-Guided Concept Bottleneck Models with Minimal Annotations

Proposes VH-CBM, combining VLM embeddings and minimal annotations via Gaussian Processes to enhance concept accuracy and calibration.

Nicola Debole, Andrea Passerini, Stefano Teso et al.

2026-05-13 39
cs.CV 2605.13018

OCH3R: Object-Centric Holistic 3D Reconstruction

OCH3R uses Transformer for single-image multi-object 3D reconstruction, achieving high speed and accuracy.

Yi Du, Yang You, Xiang Wan et al.

2026-05-13 42
cs.CV 2605.12650

CRAFT: Clinical Reward-Aligned Finetuning for Medical Image Synthesis

CRAFT uses reward-based fine-tuning with foundation models to improve clinical alignment in medical image synthesis, reducing low-quality hallucinations.

Yunsung Chung, Alex El Darzi, Carlo El Khoury et al.

2026-05-13 1 citations 55
cs.CV 2605.12501

Covering Human Action Space for Computer Use: Data Synthesis and Benchmark

CUActSpot benchmark enhances GUI complex interaction performance via data synthesis and multimodal evaluation; Phi-Ground-Any-4B excels.

Miaosen Zhang, Xiaohan Zhao, Zhihong Tan et al.

2026-05-13 209
cs.CV 2605.12500

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

SenseNova-U1 unifies multimodal understanding and generation via NEO-unify architecture, enhancing vision-language model performance.

Haiwen Diao, Penghao Wu, Hanming Deng et al.

2026-05-13 438
cs.CV 2605.12496

CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives

CausalCine achieves real-time multi-shot video generation using a causal autoregressive framework, significantly enhancing cross-shot coherence and interactivity.

Yihao Meng, Zichen Liu, Hao Ouyang et al.

2026-05-13 468
cs.CV 2605.12495

AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward

AlphaGRPO enhances UMMs' multimodal generation via Decompositional Verifiable Reward, significantly improving benchmarks like GenEval.

Runhui Huang, Jie Wu, Rui Yang et al.

2026-05-13 235
cs.CV 2605.12480

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation

OmniNFT enhances audio-video generation quality and synchronization through a modality-aware online diffusion RL framework.

Guohui Zhang, XiaoXiao Ma, Jie Huang et al.

2026-05-13 395
cs.CV 2605.12451

FuTCR: Future-Targeted Contrast and Repulsion for Continual Panoptic Segmentation

FuTCR framework improves new-class panoptic quality by up to 28% in continual panoptic segmentation while enhancing base-class performance.

Nicholas Ikechukwu, Keanu Nichols, Deepti Ghadiyaram et al.

2026-05-13 183
cs.CV 2605.12437

3D Gaussian Splatting for Efficient Retrospective Dynamic Scene Novel View Synthesis with a Standardized Benchmark

Proposes a synchronized multi-view 3D Gaussian Splatting framework for efficient dynamic scene reconstruction without temporal coupling, validated on a new benchmark.

Yunxiao Zhang, Suryansh Kumar

2026-05-13 38
cs.CV 2605.12119

MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics

MoCam employs structured denoising dynamics in diffusion models to unify geometry and appearance for robust novel view synthesis.

Haofeng Liu, Yang Zhou, Ziheng Wang et al.

2026-05-12 36
cs.CV 2605.12112

When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy

Introducing perceptual entropy to prevent diversity collapse in flow-based RLHF, achieving a score of 0.734 and diversity of 0.989, surpassing baselines.

Xiaofeng Tan, Jun Liu, Bin-Bin Gao et al.

2026-05-12 41
cs.CV 2605.12074

BARISTA: A Multi-Task Egocentric Benchmark for Compositional Visual Understanding

BARISTA employs dense scene graphs in egocentric videos for multi-task understanding, covering 185 real coffee-making videos with over 3.6 million annotations.

Patrick Knab, Orgest Xhelili, Inis Buzi et al.

2026-05-12 64
cs.CV 2605.11960

Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters

Chronicles-OCR evaluates VLLMs' cross-temporal visual perception of Chinese character evolution using a Stage-Adaptive Annotation Paradigm with 2,800 images.

Gengluo Li, Shangpin Peng, Xingyu Wan et al.

2026-05-12 6
cs.CV 2605.11594

PointForward: Feedforward Driving Reconstruction through Point-Aligned Representations

PointForward uses point-aligned sparse 3D queries for fast, multi-view consistent driving scene reconstruction, outperforming pixel-aligned methods.

Cheng Chi, Xianqi Wang, Hongcheng Luo et al.

2026-05-12 33
cs.CV 2605.11567

Dynamic Execution Commitment of Vision-Language-Action Models

A3 mechanism dynamically determines execution horizon via self-speculative prefix verification, improving success rate and inference efficiency across tasks.

Feng Chen, Xianghui Wang, Yuxuan Chen et al.

2026-05-12 36
cs.CV 2605.11559

When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs

LaSCD method reduces visual hallucinations using Laplacian energy, improving accuracy.

Fanpu Cao, Xin Zou, Xuming Hu et al.

2026-05-12 4
cs.CV 2605.10873

CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation

CADBench evaluates 11 models on 18,000 samples across 6 metrics, revealing performance gaps in multimodal 3D CAD generation.

Anna C. Doris, Jacob Thomas Sony, Ghadi Nehme et al.

2026-05-12 58
cs.CV 2605.10307

PaMoSplat: Part-Aware Motion-Guided Gaussian Splatting for Dynamic Scene Reconstruction

Proposed PaMoSplat integrates part-aware modeling and motion priors for dynamic Gaussian splatting, achieving superior rendering and tracking accuracy.

Yinan Deng, Jianyu Dou, Jiahui Wang et al.

2026-05-11 65
Prev 1 ... 29 30 31 32 33 34 35 ... 137 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home