PTQ4ARVG: Post-Training Quantization for AutoRegressive Visual Generation Models
PTQ4ARVG framework achieves 6-bit quantization for ARVG models while maintaining competitive performance.
Xuewen Liu, Zhikai Li, Jing Zhang et al.
PTQ4ARVG framework achieves 6-bit quantization for ARVG models while maintaining competitive performance.
Xuewen Liu, Zhikai Li, Jing Zhang et al.
DeepSeek-OCR 2 uses DeepEncoder V2 for visual causal flow, achieving a 3.73% performance boost on OmniDocBench v1.5.
Haoran Wei, Yaofeng Sun, Yukun Li
Youtu-Parsing achieves 5-11x speedup via high-parallelism decoding, setting SOTA performance on OmniDocBench and olmOCR-bench.
Haoyu Cao, Kun Yin, Yunfei Wu et al.
HINT introduces a hierarchical interaction autoregressive diffusion framework, achieving a low FID of 3.100 and high semantic coherence for multi-human motion generation.
Mengge Liu, Yan Di, Gu Wang et al.
VGGT-SLAM 2.0 employs a novel factor graph and attention-based verification to achieve real-time dense scene reconstruction, reducing pose error by 23%.
Dominic Maggio, Luca Carlone
GenAgent employs agentic multimodal reasoning with tool invocation, achieving +23.6% performance on GenEval++.
Kaixun Jiang, Yuzheng Wang, Junjie Zhou et al.
SRA 2 accelerates diffusion model training using VAE feature alignment, adding only 4% GFLOPs.
Mengmeng Wang, Dengyang Jiang, Liuzhuozheng Li et al.
CoT-Seg leverages chain-of-thought reasoning and self-correction, enabling training-free, robust segmentation with 10%+ improvements on challenging datasets.
Shiu-hong Kao, Chak Ho Huang, Huaiqian Liu et al.
Proposes PPIA, a black-box physical prompt injection attack using visual observation with 98% success, exploiting environment cues.
Chen Ling, Kai Hu, Hangcheng Liu et al.
ActionMesh integrates temporal diffusion and autoencoder to rapidly generate topology-consistent animated 3D meshes.
Remy Sabathier, David Novotny, Niloy J. Mitra et al.
SAMTok discretizes masks into two tokens, enabling pixel-level understanding without architectural changes.
Yikang Zhou, Tao Zhang, Dengxian Gong et al.
Proposes an iterative self-correction framework using vision-language critics to improve complex prompt image generation, achieving 16.9% accuracy gain.
Shantanu Jaiswal, Mihir Prabhudesai, Nikash Bhardwaj et al.
RayRoPE employs projective ray encoding with depth prediction and uncertainty modeling, ensuring SE(3) invariance and scene geometry adaptation, boosting novel-view synthesis by 24%.
Yu Wu, Minsik Jeon, Jen-Hao Rick Chang et al.
Proposes Soft Tail-dropping Adaptive Tokenizer (STAT), dynamically adjusting token count based on image complexity, improving reconstruction and generation quality.
Zeyuan Chen, Kai Zhang, Zhuowen Tu et al.
AVIR combines lightweight retrieval and adaptive filtering to reduce 70% pages, achieving 84.58% ANLS in multi-page VQA.
Zongmin Li, Yachuan Li, Lei Kang et al.
SeLop method uses low-rank orthogonal projection to remove spurious bias, improving face forgery detection generalization.
Chi Wang, Xinjue Hu, Boyu Wang et al.
AIGVDBench: benchmark with 31 models, 440k videos, evaluating 33 detectors, offering comprehensive analysis.
Long Ma, Zihao Xue, Yan Wang et al.
HuDA model uses human detection and temporal alignment to enhance video generation, achieving a 73% win rate.
Kumar Ashutosh, XuDong Wang, Xi Yin et al.
Molmo2 is an open-source vision-language model with pixel-level grounding, outperforming existing open and proprietary models on video understanding tasks.
Christopher Clark, Jieyu Zhang, Zixian Ma et al.
Improved video generation physics plausibility using WMReward and VJEPA-2, achieving 62.64% in ICCV 2025 challenge.
Jianhao Yuan, Xiaofeng Zhang, Felix Friedrich et al.