Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Janus-Pro scales up to 7B parameters, employing optimized training, expanded data, and decoupled visual encoding, achieving state-of-the-art multimodal understanding and text-to-image generation.
Key Findings
Methodology
Janus-Pro adopts a decoupled visual encoding architecture, separating understanding and generation tasks. It utilizes SigLIP for high-dimensional semantic feature extraction for understanding, while employing VQ tokenizer to discretize images for generation. The training process involves a three-stage pipeline with optimized strategies: extended Stage I pixel modeling, direct training on dense text descriptions in Stage II, and data ratio adjustments in Stage III. The training data is expanded with multi-source datasets like YFCC and Docmatix, supplemented with 72 million synthetic aesthetic samples to improve stability and quality. The model scales from 1.5B to 7B parameters, trained on the HAIL-LLM framework with AdamW optimizer, incorporating learning rate schedules and early stopping to ensure convergence across tasks.
Key Results
- Janus-Pro-7B achieves 79.2 on the MMBench multimodal understanding benchmark, surpassing Janus (69.4) and TokenFlow (68.9), demonstrating superior multi-task capabilities. It performs well on POPE, MME-Perception, GQA, with significant accuracy improvements.
- In text-to-image instruction-following, Janus-Pro-7B scores 80% on GenEval, outperforming Janus (61%) and DALL-E 3 (67%), indicating enhanced instruction adherence and detail fidelity. On DPG-Bench, it scores over 84, showing strong dense instruction following.
- Training data expansion and strategy optimization lead to more stable, detailed, and aesthetically pleasing images, validated across multiple benchmarks. Larger models (7B) converge faster and outperform smaller counterparts, confirming scalability and robustness.
Significance
This work addresses longstanding challenges in multi-task multimodal models by decoupling visual encoding, thus reducing task conflict and improving overall performance. It sets new state-of-the-art results in both understanding and generation, advancing the field towards more capable, stable, and scalable multimodal AI systems. The approach paves the way for practical applications in content creation, virtual assistants, and AR/VR, with broad industry implications. The combination of architecture innovation, training optimization, and data augmentation demonstrates a comprehensive strategy for future large-scale multimodal models.
Technical Contribution
The key technical contributions include the decoupled visual encoding architecture using SigLIP and VQ tokenizer, which effectively separates understanding and generation feature spaces. The optimized training pipeline—extending pixel modeling, focusing on dense text descriptions, and adjusting data ratios—significantly enhances training efficiency and model performance. The large-scale data augmentation, including synthetic aesthetic data, improves model stability and visual quality. The parameter scaling from 1.5B to 7B validates the architecture’s scalability and training effectiveness. These innovations collectively push the boundaries of multimodal AI, enabling models to better balance understanding and generation tasks.
Novelty
This work is the first to systematically apply visual encoding decoupling in a unified multimodal model, combining SigLIP and VQ tokenization to address the inherent conflict between understanding and generation representations. Unlike previous models that used a single shared encoder, Janus-Pro’s task-specific visual encoders ensure better feature disentanglement. The optimized training strategy, including extended pixel modeling and direct dense text description training, further distinguishes it from prior approaches. The integration of synthetic aesthetic data for training stability and quality is also a novel aspect, making this a comprehensive advancement in multimodal model design.
Limitations
- The input resolution remains limited to 384×384, constraining performance on fine-grained tasks such as OCR and microstructure recognition. Increasing resolution is necessary for further improvements.
- Generated images, while semantically rich, still lack fine details, especially in small regions like facial features, due to the limitations of the visual tokenizer and reconstruction loss.
- Training costs are high, particularly with 7B parameters, requiring substantial computational resources. Efficient model compression and inference acceleration are needed for broader deployment.
- The current architecture may face challenges in extreme out-of-distribution scenarios, indicating a need for better robustness and generalization strategies.
- Data quality and diversity still limit the model’s capabilities; more high-quality, diverse datasets are essential for future improvements.
Future Work
Future directions include increasing input and output resolutions to improve fine details, integrating more diverse and high-quality datasets, and exploring more efficient training and inference techniques. Additionally, enhancing model robustness and interpretability remains a priority, along with extending the architecture to handle more modalities like audio and video. The authors also plan to investigate continual learning strategies to adapt to evolving data distributions, aiming for more generalizable and resource-efficient multimodal AI systems.
AI Executive Summary
Multimodal understanding and generation have become pivotal in advancing artificial intelligence, yet existing models often struggle with balancing the two tasks, leading to conflicting representations and limited performance. Traditional approaches rely on shared visual encoders, which inadvertently cause interference between understanding and generation, resulting in suboptimal outputs and instability. As the demand for more sophisticated AI systems grows, researchers seek architectures that can effectively decouple these tasks, allowing specialized processing pathways.
Janus-Pro emerges as a groundbreaking solution, building upon the previous Janus model. Its core innovation lies in decoupling visual encoding for understanding and generation, employing SigLIP for semantic feature extraction and VQ tokenizer for image discretization. This separation ensures that each task has dedicated feature representations, reducing mutual interference. The architecture is complemented by an optimized training pipeline, which extends pixel modeling in Stage I, focuses on dense text descriptions in Stage II, and adjusts data ratios in Stage III to maximize training efficiency.
The training data is significantly expanded, incorporating multi-source datasets like YFCC and Docmatix, and adding 72 million synthetic aesthetic samples. These enhancements improve the model’s stability, aesthetic quality, and generalization ability. The model scales from 1.5 billion to 7 billion parameters, demonstrating excellent scalability and faster convergence, validated through extensive experiments.
Results across multiple benchmarks underscore Janus-Pro’s superiority. In multimodal understanding, it surpasses previous SOTA models, achieving 79.2 on MMBench. In text-to-image generation, it scores 80% on GenEval, outperforming models like DALL-E 3 and Janus. The model’s ability to follow complex instructions and produce detailed, aesthetically pleasing images marks a significant step forward.
This work’s significance extends beyond technical achievements. It provides a robust framework for future multimodal AI development, addressing core challenges of task conflict and data efficiency. Its broad applicability spans content creation, virtual assistants, AR/VR, and beyond, promising a future where AI seamlessly integrates multiple modalities for richer, more natural interactions.
Despite these advances, limitations remain. Input resolution constraints hinder fine-grained tasks, and high computational costs pose deployment challenges. Future efforts will focus on higher resolution, more diverse data, and efficiency improvements, aiming to realize the full potential of multimodal AI in real-world applications.
Deep Dive
Abstract
In this work, we introduce Janus-Pro, an advanced version of the previous work Janus. Specifically, Janus-Pro incorporates (1) an optimized training strategy, (2) expanded training data, and (3) scaling to larger model size. With these improvements, Janus-Pro achieves significant advancements in both multimodal understanding and text-to-image instruction-following capabilities, while also enhancing the stability of text-to-image generation. We hope this work will inspire further exploration in the field. Code and models are publicly available.
Cited By (20)
DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models
VIG-RL: Learning to Search and Insert for Verified Image Grounding
Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering
Transferability Between Understanding and Generation in Unified Multimodal Models
One Rewrite to Fix Them All? Type-Aware Repair Allocation for Text-to-Image Prompt Optimization
Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes
Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
Improving Sample Diversity in Autoregressive Text-to-Image Generation via Cluster Truncation
Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation
Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis
Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing
Orca: The World is in Your Mind
Unified Audio Intelligence Without Regressing on Text Intelligence
SymbOmni: Evolving Agentic Omni Models via Symbolic Concept Learning
Evaluating and Understanding Model Editing for Medical Vision Language Models
BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language
SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation
Scaling GUI Agents with Visual State Transitions
Show Me Examples: Inferring Visual Concepts from Image Sets
RADIANCE: Relative Adaptive Denoising with IP-Adapter for Novel Concept Enhancement