Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets
Step1X-3D combines VAE-DiT and diffusion models to generate high-quality, controllable 3D assets with a curated 2M dataset.
Weiyu Li, Xuanyang Zhang, Zheng Sun et al.
Step1X-3D combines VAE-DiT and diffusion models to generate high-quality, controllable 3D assets with a curated 2M dataset.
Weiyu Li, Xuanyang Zhang, Zheng Sun et al.
Seed1.5-VL combines a 532M vision encoder with a 20B MoE LLM, achieving SOTA on 38 of 60 benchmarks, excelling in multimodal reasoning.
Dong Guo, Faming Wu, Feida Zhu et al.
Proposes Sparse Concept Layer in ViT for rule extraction, improving accuracy by 5.14%.
Parth Padalkar, Gopal Gupta
Filtered LLaVA dataset using Toxic-BERT and LlavaGuard, removing 7,531 toxic image-text pairs.
Karthik Reddy Kanjula, Surya Guthikonda, Nahid Alam et al.
GPT-Image model for image restoration enhances visual quality but lacks pixel-level structural fidelity.
Hao Yang, Yan Yang, Ruikun Zhang et al.
BrickGPT uses large-scale stable datasets and physics-aware inference to generate structurally stable brick models from text prompts.
Ava Pun, Kangle Deng, Ruixuan Liu et al.
FG-CLIP leverages 1.6 billion long caption-image pairs, 12M region annotations, and 10M hard negatives to enhance fine-grained alignment.
Chunyu Xie, Bin Wang, Fanjing Kong et al.
DenseGrounding improves 3D visual grounding accuracy by 5.81% using HSSE and LSE modules to enhance visual and textual semantics.
Henry Zheng, Hao Shi, Qihang Peng et al.
Proposes a four-stage roadmap from modular to native multimodal reasoning, emphasizing the potential of N-LMRMs for scalable, autonomous AI.
Yunxin Li, Zhenyu Liu, Zitao Li et al.
VLA models integrate vision, language, and actions, enabling robots to understand and operate autonomously in complex environments.
Ranjan Sapkota, Yang Cao, Konstantinos I. Roumeliotis et al.
Unified multimodal model combining diffusion and autoregressive mechanisms, enhancing understanding and generation capabilities.
Shanshan Zhao, Xinjie Zhang, Jintao Guo et al.
DetoxAI is a post-hoc debiasing toolkit for deep vision models, supporting state-of-the-art algorithms and fairness metrics, enhancing model fairness without retraining.
Ignacy Stępka, Lukasz Sztukiewicz, Michał Wiliński et al.
SpatialLLM integrates 3D-aware data and multi-stage training, achieving 8.7% better than GPT-4o in 3D spatial reasoning tasks.
Wufei Ma, Luoxin Ye, Celso M de Melo et al.
JointDiT employs diffusion transformers for RGB-depth joint modeling, enabling high-fidelity joint generation and accurate depth estimation.
Kwon Byung-Ki, Qi Dai, Lee Hyoseok et al.
Common3D employs self-supervised learning to build category-agnostic 3D deformable models from videos, improving pose and correspondence estimation.
Leonhard Sommer, Olaf Dünkel, Christian Theobalt et al.
CMT: Boundary-representation-based multimodal cascade MAR, improves coverage by 10.68%, supports complex CAD generation.
Jianyu Wu, Yizhou Wang, Xiangyu Yue et al.
RepText employs a copying mechanism within a diffusion framework to accurately replicate multilingual visual text without understanding, achieving 85.4% accuracy on ICDAR.
Haofan Wang, Yujia Xu, Yimeng Li et al.
The 4th Monocular Depth Estimation Challenge improved 3D F-Score to 23.05% using affine-invariant predictions.
Anton Obukhov, Matteo Poggi, Fabio Tosi et al.
Proposes MR. Video based on MapReduce, achieving 10%+ accuracy boost on LVBench for long video understanding.
Ziqi Pang, Yu-Xiong Wang
Eagle 2.5 employs Automatic Degrade Sampling and Image Area Preservation to enhance long-video understanding, achieving 72.4% on Video-MME with 8B parameters, comparable to GPT-4o.
Guo Chen, Zhiqi Li, Shihao Wang et al.