Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching
DM-Align jointly distills and aligns video generators, reaching 84.40 VBench at only 4 NFE.
Jiuzhou Lin, Junlong Wu, Fei Zuo et al.
DM-Align jointly distills and aligns video generators, reaching 84.40 VBench at only 4 NFE.
Jiuzhou Lin, Junlong Wu, Fei Zuo et al.
Reflection-Aware GRPO integrates diffusion reflection and counterfactual path synthesis to enhance semantic fidelity and realism in visual generation.
Junlong Wu, Jiuzhou Lin, Jia Sun et al.
RoboTok learns a latent 3D hand trajectory space for internet video retrieval, boosting robot manipulation demonstration matching.
Howard Qian, Yiting Chen, Yunfei Xie et al.
SolarWM employs a reconfigurable multi-source data engine and backbone-native adaptation to train long-horizon video world models, enabling real-time interaction.
Junchao Huang, Guian Fang, Shengju Qian et al.
Introduces RIG-BENCH, a comprehensive benchmark evaluating reasoning-driven image generation across four domains, revealing a significant gap between current models and human-level reasoning.
Yutong Liu, Nan Huang, Xu Cao et al.
SR-Edit achieves region-aware image editing via self-refinement, significantly improving non-edit region preservation.
Andong Wang, Zehua Chen, Yuxuan Jiang et al.
Introduced VMetaphor-Bench to benchmark visual metaphor generation in T2I models, revealing challenges in cross-domain mapping and compositional structuring.
Chuer Chen, Zichen Wang, Yi He et al.
VIPS uses pseudo-simulation with real data to evaluate V2I autonomous driving robustness efficiently.
Hoonhee Cho, Jae-Young Kang, Giwon Lee et al.
CA-OPD uses teacher confidence to repair student rollouts, raising ScreenSpot-Pro by 9.50 points and OCRBench-v2 English by 6.72 points.
Menghao Li, Linjie Mu, Yin Wang et al.
DPA method decouples product-agnostic anomaly representations for zero-shot anomaly generation, enhancing detection performance.
Hang Yao, Yansheng Fu, Ming Liu et al.
InstEditSeg uses instruction-driven image editing for polyp and skin lesion segmentation, significantly enhancing cross-domain generalization.
Ziquan Liu, Zhewei Zhu, Xuyang Shi
ZipTok3D achieves high-fidelity 3D reconstruction with compact token prefixes, using only one token on ShapeNet.
Mingda Lin, Weijie Wang, Zeyu Zhang et al.
SpatialGuard employs structured layout and verification to enhance spatial fidelity in complex 3D text-to-image generation, achieving state-of-the-art results.
Ziyun Qian, Zizhi Chen, Yizhou Liu et al.
H3-World leverages MiniMax-H3 to enable precise, temporally grounded world control via natural language, with minimal fine-tuning and high generalization.
Danze Chen, Zeqing Wang, Ziyue Lin et al.
Layer-wise probing of V-JEPA 2 and VideoMAE-v2 reveals that camera motion is encoded mainly in mid-layers, forming smooth trajectories in feature space, with spline interpolation improving motion coherence.
Sharon S. Musa, Fereshteh Forghani, Harrish Thasarathan et al.
Gekko uses relative reconstruction error improvement to enhance 3D features without 3D labels, outperforming CroCo.
Thibaut Loiseau, Guillaume Bourmaud, Vincent Lepetit
TempCloze evaluates visual temporal reasoning by identifying the missing middle segment from video clips, revealing that alignment remains the main bottleneck.
Wenqi Pei, Henry Hengyuan Zhao, Yilai Liu et al.
CameraEditor uses video diffusion models for camera parameter control, enhancing geometric precision.
Xin Shen, Chengyou Jia, Keshuo Xing et al.
One prompt enables watermark laundering through foundation image models, significantly reducing watermark reliability.
Jidong Yang, Qi Li, Wei Zong et al.
SciGram framework generates 1.4M visual instructions, advancing scientific diagram understanding and surpassing SOTA.
Raul Ortega, José Manuel Gómez-Pérez