HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

TL;DR

HunyuanOCR-1.5 enhances OCR performance using DFlash acceleration and Agentic Data Flow, achieving 6.37x inference speedup.

cs.CV 🔴 Advanced 2026-07-06 23 views
Gengluo Li Xingyu Wan Shangpin Peng Weinong Wang Hao Feng Yongkun Du Binghong Wu Zheng Ruan Zhiqiong Lu Liang Wu Pengyuan Lyu Huawen Shen Zibin Lin Shijing Hu Jieneng Yang Hongbing Wen Guanghua Yu Hong Liu Bochao Wang Can Ma Han Hu Chengquan Zhang Yu Zhou
OCR lightweight vision-language model inference acceleration data flow

Key Findings

Methodology

HunyuanOCR-1.5 employs DFlash for inference acceleration and Agentic Data Flow for data construction. DFlash uses a block-diffusion draft model to predict multiple candidate tokens in parallel, significantly improving decoding efficiency for long outputs. Agentic Data Flow automates material search and quality verification, targeting long-tail tasks like low-resource multilingual parsing and ancient-script OCR.

Key Results

  • HunyuanOCR-1.5 achieves a 6.37x speedup in Transformer inference and a 2.14x speedup under vLLM, making it the fastest among lightweight OCR VLMs.
  • It ranks top-tier in end-to-end OCR solutions on OmniDocBench v1.6, setting new performance milestones in long-tail tasks like ancient-script OCR and low-resource multilingual parsing.
  • With upgraded pretraining and post-training strategies, HunyuanOCR-1.5 extends its capability in high-resolution, long-context, and multi-task scenarios.

Significance

HunyuanOCR-1.5 holds significant implications for academia and industry. It addresses the inefficiency of long-output decoding and enhances model capabilities in long-tail tasks through Agentic Data Flow. Its lightweight design offers deployment advantages, particularly in applications requiring efficient and accurate OCR.

Technical Contribution

Technically, HunyuanOCR-1.5 achieves substantial acceleration and capability enhancement through DFlash and Agentic Data Flow. Unlike existing SOTA methods, it improves OCR task coverage and inference speed without altering the architecture, through data and training strategy upgrades.

Novelty

HunyuanOCR-1.5 is the first to apply DFlash to OCR decoding, accelerating long-output generation. Agentic Data Flow significantly enhances performance in long-tail tasks through an automated data construction system, offering innovation over traditional manual data collection methods.

Limitations

  • In extremely complex document parsing scenarios, the model may still face performance bottlenecks, especially with multi-page or multi-image inputs.
  • For some low-resource languages, performance may be limited by the quality and coverage of training data.

Future Work

Future work could focus on further optimizing DFlash's parallel decoding capabilities and expanding Agentic Data Flow to support more types of long-tail tasks. Exploring lightweight deployment on mobile devices is also a key direction.

AI Executive Summary

HunyuanOCR-1.5 significantly enhances OCR task efficiency and capability through innovative DFlash inference acceleration and Agentic Data Flow data construction systems. Existing OCR solutions often face decoding delay issues when handling long outputs, but HunyuanOCR-1.5 achieves a 6.37x inference speedup using DFlash, making it the fastest among lightweight OCR VLMs.

In terms of capability enhancement, the Agentic Data Flow system significantly improves model performance in long-tail tasks such as low-resource multilingual parsing, ancient-script recognition, and multi-image QA through automated material search, quality verification, and data pipeline development. This system transforms model weaknesses into executable data requirements, driving a closed-loop data production process.

HunyuanOCR-1.5 performs exceptionally well on OmniDocBench v1.6, achieving new performance milestones in long-tail tasks. Combined with upgraded pretraining and post-training strategies, the model extends its capability boundary in high-resolution, long-context, and multi-task scenarios. The model weights and training code will be released to promote research and applications in the OCR field.

Deep Analysis

Background

Optical Character Recognition (OCR) technology is foundational for digitizing information, traditionally used for simple text transcription. However, as the demand for machine intelligence grows, OCR needs to support more complex text-centric visual tasks such as document parsing, information extraction, and visual question answering. Existing OCR-specialized vision-language models (VLMs) are often designed for document parsing, struggling to meet the diverse OCR needs of real-world applications.

Core Problem

Traditional cascaded pipelines often face error propagation and architectural redundancy issues when handling complex OCR tasks. Existing OCR-specialized VLMs mainly focus on document parsing, struggling with tasks like text spotting in open scenarios, structured field extraction, multilingual text-image translation, and reasoning across multiple pages or images. A true OCR-specialized model should cover diverse OCR tasks, not just document parsing.

Innovation

HunyuanOCR-1.5 introduces significant innovations through DFlash inference acceleration and Agentic Data Flow data construction systems. DFlash uses a block-diffusion draft model to predict multiple candidate tokens in parallel, significantly improving decoding efficiency for long outputs. Agentic Data Flow automates material search and quality verification, targeting long-tail tasks like low-resource multilingual parsing and ancient-script OCR.

Methodology

  • �� Use DFlash's block-diffusion draft model for parallel candidate token prediction, improving decoding efficiency.
  • �� Employ Agentic Data Flow system for automated material search, quality verification, and data pipeline development, enhancing model performance in long-tail tasks.
  • �� Upgrade pretraining and post-training strategies to extend model capability in high-resolution, long-context, and multi-task scenarios.

Experiments

Experiments were conducted on OmniDocBench v1.6, achieving a 6.37x inference speedup using DFlash and a 2.14x speedup under vLLM. The Agentic Data Flow system enabled the model to set new performance milestones in long-tail tasks like low-resource multilingual parsing, ancient-script recognition, and multi-image QA.

Results

HunyuanOCR-1.5 achieves a 6.37x speedup in Transformer inference and a 2.14x speedup under vLLM, making it the fastest among lightweight OCR VLMs. It ranks top-tier in end-to-end OCR solutions on OmniDocBench v1.6, setting new performance milestones in long-tail tasks like ancient-script OCR and low-resource multilingual parsing.

Applications

HunyuanOCR-1.5 is suitable for real-world deployment scenarios requiring efficient and accurate OCR, especially in complex document parsing, multilingual text-image translation, and multi-image QA tasks. Its lightweight design offers advantages in mobile devices and low-computation environments.

Limitations & Outlook

Despite significant advances in inference speed and capability coverage, HunyuanOCR-1.5 may still face performance bottlenecks in extremely complex document parsing scenarios. Additionally, performance for some low-resource languages may be limited by the quality and coverage of training data. Future improvements could focus on further optimizing DFlash's parallel decoding capabilities and expanding Agentic Data Flow to support more types of long-tail tasks.

Plain Language Accessible to non-experts

Imagine you're in a library, and HunyuanOCR-1.5 is like a super librarian. It not only finds the books you need quickly but also organizes all the books on the shelves, whether they're ancient manuscripts or modern magazines. Traditional librarians might take a long time to complete these tasks, but HunyuanOCR-1.5 uses a special method called DFlash to handle multiple tasks simultaneously, greatly speeding things up. It also has a helper called Agentic Data Flow that automatically collects and organizes various materials in the library, ensuring all books are correctly identified and categorized.

ELI14 Explained like you're 14

Hey, imagine you're playing a super cool game called HunyuanOCR-1.5! The goal of the game is to quickly find and recognize all sorts of hidden texts, like hunting for treasure in a giant maze. You have a super helper called DFlash that helps you handle multiple tasks at once, letting you speed through the game! Plus, you have a smart assistant called Agentic Data Flow that automatically gathers all the clues and tools you need in the game, making it easy to win! Isn't that awesome?

Glossary

DFlash (Block Diffusion)

DFlash is an inference acceleration technique for speeding up long-output generation by predicting multiple candidate tokens in parallel.

Used in HunyuanOCR-1.5 to improve decoding efficiency for long outputs.

Agentic Data Flow

Agentic Data Flow is an automated data construction system that enhances model performance in long-tail tasks through material search and quality verification.

Used in HunyuanOCR-1.5 for data construction, especially in low-resource multilingual parsing and ancient-script OCR.

OmniDocBench v1.6

OmniDocBench v1.6 is a benchmark platform for evaluating OCR model performance.

HunyuanOCR-1.5 performed exceptionally well on this platform, especially in long-tail tasks.

Transformer

Transformer is a neural network architecture used in natural language processing and computer vision, known for its parallel processing capabilities.

Used in HunyuanOCR-1.5 for inference acceleration.

vLLM

vLLM is a framework for efficient large-scale language model inference, supporting high-performance model deployment.

HunyuanOCR-1.5 achieved a 2.14x inference speedup under vLLM.

Open Questions Unanswered questions from this research

  • 1 How can DFlash's parallel decoding capabilities be further optimized to support more complex document parsing tasks?
  • 2 Can Agentic Data Flow be expanded to other types of long-tail tasks, such as video subtitle extraction?

Applications

Immediate Applications

Document Parsing

HunyuanOCR-1.5 can be used for rapid parsing of complex document structures, suitable for enterprise document management and digital archiving.

Long-term Vision

Multilingual Translation

By expanding Agentic Data Flow, HunyuanOCR-1.5 has the potential to support real-time multilingual translation, facilitating cross-cultural communication.

Abstract

We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies document parsing, text spotting, information extraction, text-image translation, and multi-image document understanding within a single end-to-end VLM. Building upon the lightweight architecture of HunyuanOCR-1.0, HunyuanOCR-1.5 does not redesign the backbone, but systematically improves both efficiency and capability. For efficiency, we adapt DFlash to OCR decoding, significantly reducing the latency of long structured outputs such as dense documents, tables, and formulas while preserving output distribution. Powered by DFlash, HunyuanOCR-1.5 achieves a 6.37x Transformer inference speedup and a 2.14x speedup under vLLM, delivering the fastest inference among lightweight OCR VLMs. For capability, we propose Agentic Data Flow, an agent-driven data construction system that transforms model weaknesses into executable data requirements and autonomously performs material search, quality verification, and pipeline development. It substantially improves long-tail capabilities in ancient-script OCR, fine-grained chart and table parsing, multi-image text-centric QA, low-resource multilingual parsing, and document hallucination evaluation. HunyuanOCR-1.5 ranks among the top-tier end-to-end OCR solutions on OmniDocBench v1.6 while achieving new performance milestones across these long-tail tasks. Combined with an upgraded pretraining and post-training recipe, HunyuanOCR-1.5 further extends its capability in high-resolution, long-context, and multi-task scenarios. Experiments demonstrate faster inference, broader OCR capability coverage, and the deployment advantages of a lightweight end-to-end model. We will release the model weights and training code to support future research and real-world OCR applications.

cs.CV