UFOGen: You Forward Once Large Scale Text-to-Image Generation via Diffusion GANs
UFOGen achieves ultra-fast one-step text-to-image generation via Diffusion GANs, significantly reducing computational costs.
Yanwu Xu, Yang Zhao, Zhisheng Xiao et al.
UFOGen achieves ultra-fast one-step text-to-image generation via Diffusion GANs, significantly reducing computational costs.
Yanwu Xu, Yang Zhao, Zhisheng Xiao et al.
GPT-4V-based VLM demonstrates superior scene understanding and causal reasoning in autonomous driving, outperforming existing systems by 10-15% in key metrics.
Licheng Wen, Xuemeng Yang, Daocheng Fu et al.
This paper introduces LRM, a transformer-based large-scale single-image 3D reconstruction model trained on ~1 million objects, capable of generating high-fidelity models in under 5 seconds.
Yicong Hong, Kai Zhang, Jiuxiang Gu et al.
AnyText employs diffusion models with OCR-guided multi-modal embeddings to achieve high-fidelity multilingual visual text generation and editing, outperforming baselines on 3M image-text pairs.
Yuxiang Tuo, Wangmeng Xiang, Jun-Yan He et al.
GPT-4V excels in multi-modal anomaly detection, particularly in zero/one-shot scenarios.
Yunkang Cao, Xiaohao Xu, Chen Sun et al.
FETV benchmark introduces multi-dimensional, temporal-aware fine-grained evaluation for open-source T2V models, revealing poor correlation of existing metrics with human judgment.
Yuanxin Liu, Lei Li, Shuhuai Ren et al.
Copilot4D combines VQVAE and discrete diffusion for point-cloud forecasting, cutting Chamfer distance by over 65% at 1 s and 50% at 3 s.
Lunjun Zhang, Yuwen Xiong, Ze Yang et al.
SimMMDG splits features into shared and specific parts, using contrastive learning and translation for robust multi-modal domain generalization, outperforming baselines on EPIC-Kitchens and HAC.
Hao Dong, Ismail Nejjar, Han Sun et al.
Proposes high-quality open-source diffusion models for T2V (1024×576) and I2V with content preservation, advancing video synthesis.
Haoxin Chen, Menghan Xia, Yingqing He et al.
ProtoConcepts extends prototype networks with multiple visualizations, using geometric prototype balls, achieving comparable accuracy and significantly improved interpretability.
Chiyu Ma, Brandon Zhao, Chaofan Chen et al.
Proposed Spatial-Temporal Diversification Network (STDN) enhances video domain generalization via space-time content diversity, using spatial grouping and multi-scale relation modeling.
Kun-Yu Lin, Jia-Run Du, Yipeng Gao et al.
Introduced PO3D-VQA model and Super-CLEVR-3D dataset, enhancing 3D visual question answering performance.
Xingrui Wang, Wufei Ma, Zhuowan Li et al.
AntifakePrompt leverages prompt tuning on InstructBLIP, boosting deepfake detection accuracy from 71.06% to 92.11% on unseen models.
You-Ming Chang, Chen Yeh, Wei-Chen Chiu et al.
DreamCraft3D employs hierarchical generation guided by view-dependent diffusion priors, achieving high-fidelity 3D objects with mutual geometric and texture consistency.
Jingxiang Sun, Bo Zhang, Ruizhi Shao et al.
Proposes a multi-conditional diffusion model for text-guided 3D scene synthesis, outperforming state-of-the-art benchmarks.
An Vuong, Minh Nhat Vu, Toan Tien Nguyen et al.
Zero123++ employs diffusion models with multi-view joint modeling, achieving high-quality, consistent multi-view 3D generation from a single image.
Ruoxi Shi, Hansheng Chen, Zhuoyang Zhang et al.
Wonder3D employs cross-domain diffusion to efficiently generate detailed textured meshes from a single image.
Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin et al.
HallusionBench is a diagnostic benchmark for entangled language hallucination and visual illusion in large vision-language models, evaluating their reasoning robustness.
Tianrui Guan, Fuxiao Liu, Xiyang Wu et al.
Cutie adds object-level memory reading to VOS, reaching 64.0 J&F on MOSE versus 56.3 for XMem.
Ho Kei Cheng, Seoung Wug Oh, Brian Price et al.
DynamiCrafter combines video-diffusion motion priors with dual image conditioning, but the supplied text reports no numerical metrics.
Jinbo Xing, Menghan Xia, Yong Zhang et al.