Few-step Flow for 3D Generation via Marginal-Data Transport Distillation
Proposes MDT-dist, a few-step flow distillation method reducing sampling from 25 to 1-2 steps, enabling real-time 3D generation.
Zanwei Zhou, Taoran Yi, Jiemin Fang et al.
Proposes MDT-dist, a few-step flow distillation method reducing sampling from 25 to 1-2 steps, enabling real-time 3D generation.
Zanwei Zhou, Taoran Yi, Jiemin Fang et al.
OneCAT is a decoder-only autoregressive model for unified multimodal understanding, generation, and editing, eliminating external encoders for efficiency.
Han Li, Xinyu Peng, Yaoming Wang et al.
Kwai Keye-VL 1.5 employs Slow-Fast video encoding, extends context to 128K tokens, combined with multi-stage pretraining and post-training, greatly enhancing video understanding.
Biao Yang, Bin Wen, Boyang Ding et al.
FantasyHSI enables 4D human synthesis in any scene using a graph-based multi-agent framework, significantly improving task completion.
Lingzhou Mu, Qiang Wang, Fan Jiang et al.
POINTS-Reader employs a two-stage, distillation-free framework using synthetic data and self-improvement, achieving state-of-the-art document conversion performance.
Yuan Liu, Zhongyin Zhao, Le Tian et al.
GPSToken employs Gaussian parameterization for spatially adaptive image tokenization, achieving state-of-the-art FID=1.50 in image generation with 128 tokens.
Zhengqiang Zhang, Rongyuan Wu, Lingchen Sun et al.
VoxHammer is a training-free 3D editing method using inverse trajectory prediction and feature replacement.
Lin Li, Zehuan Huang, Haoran Feng et al.
Drawing2CAD employs sequence-to-sequence learning to convert SVG vector drawings into parametric CAD operation sequences.
Feiwei Qin, Shichao Lu, Junhao Hou et al.
JCo-MVTON leverages multi-modal diffusion transformers for mask-free virtual try-on, achieving SOTA performance on DressCode dataset.
Aowen Wang, Wei Li, Hao Luo et al.
UniGen framework uses CoMoE module and WeaveNet mechanism for efficient controllable image generation, excelling on Subjects-200K and MultiGen-20M datasets.
Guoqing Zhang, Xingtong Ge, Lu Shi et al.
MV-RAG integrates retrieval with multiview diffusion, significantly improving out-of-domain concept 3D generation with 15%+ performance boost.
Yosef Dayani, Omer Benishu, Sagie Benaim
StreamMem employs query-agnostic KV cache compression, enabling efficient long video understanding with fixed memory, outperforming state-of-the-art methods.
Yanlai Yang, Zhuokai Zhao, Satya Narayan Shukla et al.
GeoSAM2 achieves 84.06% mIoU on PartObjaverse-Tiny and 74.42% on PartNetE using multi-view 2D mask prediction for 3D segmentation.
Ken Deng, Yunhan Yang, Jingxiang Sun et al.
PhysGM predicts 3D Gaussian and physical parameters from a single image, enabling real-time 4D scene synthesis in under one minute.
Chunji Lv, Zequn Chen, Donglin Di et al.
EgoTwin uses a diffusion transformer to generate consistent first-person video and human motion.
Jingqiao Xiu, Fangzhou Hong, Yicong Li et al.
DyCrowd framework achieves spatio-temporal consistent 3D reconstruction of dynamic crowds from large-scene videos.
Hao Wen, Hongbo Kang, Jian Ma et al.
TiP4GEN employs a dual-branch diffusion framework with 3D Gaussian Splatting for high-quality dynamic panoramic 4D scene synthesis.
Ke Xing, Hanwen Liang, Dejia Xu et al.
EgoLoc employs hand 3D dynamics and vision-language models for zero-shot temporal interaction localization, precisely identifying contact and separation moments.
Junyi Ma, Erhang Zhang, Yin-Dong Zheng et al.
Thyme integrates code-driven image operations and reasoning, boosting multimodal model performance.
Yi-Fan Zhang, Xingyu Lu, Shukang Yin et al.
Proposes TTF-VLA, a training-free temporal fusion method combining pixel difference and attention, boosting robotic manipulation success by 4-8%.
Chenghao Liu, Jiachen Zhang, Chengxuan Li et al.