GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.CV 2608.24782

Image Difference Quantification Using Autoencoder-Based Latent Representations

Autoencoder-based latent cosine similarity quantifies image differences, achieving 98.4% class separation.

Manish Sharma, Timothy Yim, Clifton Forlines

2026-08-26 64
cs.CV 2608.24771

Ensemble of Convolutional Neural Networks for StrokePrediction: Towards Improved Diagnostic Accuracy

Ensemble CNN achieves 99.52% accuracy in stroke prediction, outperforming individual models.

Md Shahriar Sajid

2026-08-26 62
cs.CV 2608.24563

X-MULTI: VLM-based Imaging Factor Disentanglement for Factor-Aware Image Synthesis

X-MULTI leverages pretrained VLM to improve factor disentanglement in image synthesis, boosting novel combination accuracy by 12%.

Sonali Godavarthy, Matthias Neuwirth-Trapp, Tim-Felix Faasch et al.

2026-08-25 56
cs.CV 2608.23943

Luce: Relightable Gaussians for 3D Asset Generation

Luce uses Gaussian clouds for single-image to 3D asset generation, improving FID by 28%.

Mayank Singh, Michele Stoppa, Alvise Memo et al.

2026-08-25 29
cs.CV 2608.23927

GlanceWAM: Sparse Test-Time Imagination for World-Action Models

GlanceWAM achieves 72.2% success by asynchronously imagining future frames in latent space, breaking speed-success trade-off.

Linhan Wang, Zijian An, Mingyuan Zhang et al.

2026-08-25 12
cs.CV 2608.23746

CRISP: Calibration-Aware Visual State Space Duality for Remote Sensing Semantic Segmentation

CRISP enhances VSSD for remote sensing segmentation via DCO and OMP, achieving ~30M parameters with significant accuracy gains.

Kangning Wang, Haopeng Zhang, Zhiguo Jiang

2026-08-25 41
cs.CV 2608.23563

EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings

Expert-grounded distillation creates an 8B-parameter vision-language model for scalable road safety audits, outperforming larger teachers and proprietary models with 81% accuracy.

Md Thamed Bin Zaman Chowdhury, Moazzem Hossain

2026-08-25 71
cs.CV 2608.23664

Scaling Reinforcement Learning for Diffusion Models via Velocity Matching

RVM fine-tunes diffusion velocity fields without trajectories, matching or exceeding policy-gradient baselines at lower cost; no numerical scores are provided.

Jaemoo Choi, Wei Guo, Yuchen Zhu et al.

2026-08-25 17
cs.CV 2608.23503

Action-Aligned Retrieval with Pairwise Multimodal Reranking for Text-Based Person Anomaly Search

ActPair framework combines action-aligned retrieval and pairwise multimodal reranking, significantly improving text-based person anomaly search accuracy.

Thanh-Khoi Nguyen, Thanh-Nhan Vo, Trong-Thuan Nguyen et al.

2026-08-25 67
cs.CV 2608.23486

GeoWAM: Visual Geometry World Action Models for Autonomous Driving

GeoWAM predicts future scene geometry with point clouds, significantly improving autonomous driving trajectory planning.

Yiren Lu, Xin Ye, Jiaming Liu et al.

2026-08-25 87
cs.CV 2608.23405

MomADv2: Reliable Temporal Memory for End-to-End Autonomous Driving

MomADv2 employs selective state-space memory and flow-matching residual correction, significantly improving long-horizon autonomous driving stability and safety.

Ziying Song, Shengkai Zhang, Lin Liu et al.

2026-08-24 76
cs.CV 2608.23383

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

JoyAI-Echo-1.5 enhances long-form audio-visual generation with memory and geometric control, improving consistency and visual quality.

Nan Duan, Haoyang Huang, Weiyang Jin et al.

2026-08-24 1
cs.CV 2608.23189

EchoWM: Open and Enterable Omnimodal World Models

EchoWM generates 720p video, environmental sound, music, and speech, supporting continuous navigation.

Songchun Zhang, Yaowei Li, Junhao Zhuang et al.

2026-08-24 2
cs.CV 2608.22906

AquaFlow: A Monocular Gaussian Splatting SLAM for Underwater Streaming Reconstruction

AquaFlow achieves efficient underwater streaming reconstruction with monocular Gaussian Splatting SLAM, reducing localization error by 13.2% and improving PSNR by 4.74 dB.

Yingxiang Xu, Kerui Ren, Wenqi Guo et al.

2026-08-24 27
cs.CV 2608.22272

GAN-Diff : Coupling Pretrained WGAN-GP Features with Conditional Diffusion U-Nets

GAN-Diff combines pretrained WGAN-GP features with conditional diffusion U-Nets, achieving PSNR gains of 4.40dB for denoising and 3.70dB for super-resolution.

Saif Ahmed, Asadullah Hil Galib, S. M. Riaz Rahman Antu et al.

2026-08-23 36
cs.CV 2608.21972

Improved denoising diffusion probabilistic models with efficient non-diagonal covariance modeling

Proposes K-DCT covariance model to accelerate DDPM sampling, capturing non-diagonal correlations in natural images, improving quality with fewer steps.

Rui Xia, Ayan Das, Artem Artemev et al.

2026-08-22 33
cs.CV 2608.21869

GuardPaint:SpeculativeSafetyDecodingforText-to-ImageGeneration

GuardPaint uses speculative decoding to repair unsafe regions during diffusion, reducing attack success rates while preserving image quality.

Shreyash Dhoot, Paras Dhiman, Arsh Abbas Naqvi et al.

2026-08-22 24
cs.CV 2608.21784

DefaultShift: Auditing Semantic Default Shift in Accelerated Text-to-Image Models

DefaultShift quantifies semantic default shift in accelerated text-to-image models, reducing human-measured shift by 10.3%-35.1%.

Xuanhua Yin, Chuanzhi Xu, Shunqi Mao et al.

2026-08-22 2
cs.CV 2608.21360

OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs

Introduces OmniAssistBench, a dataset built via reverse engineering of internet videos, to evaluate Omni-LLMs in multi-turn multimodal interactions.

Xianyun Sun, Chaoyou Fu, Zhengye Zhang et al.

2026-08-22 70
cs.CV 2608.21305

Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning

Re3Cap employs retrieval-guided reasoning with k-core analysis to boost image captioning, achieving 8.64% improvement in relation reasoning on COCO-LN500.

Haonan Jia, Shichao Dong, Zenghui Sun et al.

2026-08-22 74
Prev 1 ... 3 4 5 6 7 8 9 ... 134 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home