RePainter: Empowering E-commerce Object Removal via Spatial-matting Reinforcement Learning
Repainter uses spatial-matting reinforcement learning to significantly enhance e-commerce image removal.
Zipeng Guo, Lichen Ma, Xiaolong Fu et al.
Repainter uses spatial-matting reinforcement learning to significantly enhance e-commerce image removal.
Zipeng Guo, Lichen Ma, Xiaolong Fu et al.
MV-Performer employs depth-guided video diffusion to synthesize 360° synchronized multi-view human videos from monocular input.
Yihao Zhi, Chenghong Li, Hongjie Liao et al.
SIGMA-GEN enables single-pass multi-subject identity-preserving image generation, achieving state-of-the-art performance using the SIGMA-SET27K dataset.
Oindrila Saha, Vojtech Krs, Radomir Mech et al.
Introduced LUR and RLUR methods to enhance uncertainty detection efficiency in driver action recognition.
Koen Vellenga, H. Joe Steinhauer, Jonas Andersson et al.
TAG amplifies tangential score components to steer diffusion sampling toward high-density regions, reducing hallucinations and improving semantic fidelity.
Hyunmin Cho, Donghoon Ahn, Susung Hong et al.
A.I.R. introduces training-free adaptive iterative reasoning for frame selection, boosting VideoQA accuracy and efficiency.
Yuanhao Zou, Shengji Jin, Andong Deng et al.
ChronoEdit leverages pretrained video generative models with explicit temporal reasoning to enhance physical consistency in image editing, achieving a 4.42/5 score on PBench-Edit.
Jay Zhangjie Wu, Xuanchi Ren, Tianchang Shen et al.
Proposes Video-in-the-Loop (ViTL), a two-stage long-video QA framework combining localization and span-aware answering, achieving up to 8.6% improvement with 50% less input.
Chendong Wang, Donglin Bai, Yifan Yang et al.
Introduced HVGC framework and BridgeDiT model for text-to-sounding video generation, outperforming existing methods.
Kaisi Guan, Xihua Wang, Zhengfeng Lai et al.
FSFSplatter achieves high-precision surface reconstruction with sparse views in 2 minutes, reducing error by 28.39%.
Yibin Zhao, Yihan Pan, Jun Nan et al.
Self-Forcing++ extends high-quality video generation to 4 minutes by sampling and distribution matching, outperforming baseline models by over 50x in length with improved fidelity.
Justin Cui, Jie Wu, Ming Li et al.
Introduces PhraseStereo, a stereo dataset for open-vocabulary phrase segmentation leveraging depth cues, with SSIM=0.601 and LPIPS=0.352 at optimal scale.
Thomas Campagnolo, Ezio Malis, Philippe Martinet et al.
LAKAN integrates facial landmarks to dynamically modulate Kolmogorov-Arnold networks, boosting deepfake detection accuracy.
Jiayao Jiang, Bin Liu, Qi Chu et al.
YOLO26 achieves end-to-end NMS-free detection, removing DFL, with ProgLoss, STAL, and MuSGD, boosting speed and accuracy for edge deployment.
Ranjan Sapkota, Rahul Harsha Cheppally, Ajay Sharda et al.
SPLICE benchmark reveals significant gaps in VLM visual reasoning compared to human performance.
Mohamad Ballout, Okajevo Wilfred, Seyedalireza Yaghoubi et al.
Introduced NeMo task for long video understanding; automated data pipeline; built NeMoBench with 31,378 QA pairs from 13,486 videos.
Zi-Yuan Hu, Shuo Liang, Duo Zheng et al.
Latent Visual Reasoning (LVR) enables end-to-end reasoning in visual embedding space, boosting perception-intensive visual question answering by 5% (71.67% vs 66.67%).
Bangzheng Li, Ximeng Sun, Jiang Liu et al.
MoReact uses diffusion models to generate reactive motions based on textual descriptions, enhancing interaction realism.
Xiyan Xu, Sirui Xu, Yu-Xiong Wang et al.
Introduces ReWatch dataset and multi-stage synthesis with Multi-Agent ReAct, boosting large vision-language models' video reasoning accuracy by over 3%.
Congzhi Zhang, Zhibin Wang, Yinchao Ma et al.
LongLive uses a causal autoregressive model with KV recaching, achieving 20.7FPS for real-time long video generation up to 240 seconds.
Shuai Yang, Wei Huang, Ruihang Chu et al.