Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought

TL;DR

Hunyuan-TurboS combines Mamba-Transformer with adaptive chain-of-thought, 560B parameters, achieving top-tier performance.

cs.CL 🔴 Advanced 2025-05-21 29 views
Tencent Hunyuan Team Ao Liu Botong Zhou Can Xu Chayse Zhou ChenChen Zhang Chengcheng Xu Chenhao Wang Decheng Wu Dengpeng Wu Dian Jiao Dong Du Dong Wang Feng Zhang Fengzong Lian Guanghui Xu Guanwei Zhang Hai Wang Haipeng Luo Han Hu Huilin Xu Jiajia Wu Jianchen Zhu Jianfeng Yan Jiaqi Zhu Jihong Zhang Jinbao Xue Jun Xia Junqiang Zheng Kai Liu Kai Zhang Kai Zheng Kejiao Li Keyao Wang Lan Jiang Lixin Liu Lulu Wu Mengyuan Huang Peijie Yu Peiqi Wang Qian Wang Qianbiao Xiang Qibin Liu Qingfeng Sun Richard Guo Ruobing Xie Saiyong Yang Shaohua Chen Shihui Hu Shuai Li Shuaipeng Li Shuang Chen Suncong Zheng Tao Yang Tian Zhang Tinghao Yu Weidong Han Weijie Liu Weijin Zhou Weikang Wang Wesleye Chen Xiao Feng Xiaoqin Ren Xingwu Sun Xiong Kuang Xuemeng Huang Xun Cao Yanfeng Chen Yang Du Zhen Yang Yangyu Tao Yaping Deng Yi Shen Yigeng Hong Yiqi Chen Yiqing Huang Yuchi Deng Yue Mao Yulong Wang Yuyuan Zeng Zenan Xu Zhanhui Kang Zhe Zhao ZhenXiang Yan Zheng Fang Zhichao Hu Zhongzhi Chen Zhuoyu Li Zongwei Li Alex Yan Ande Liang Baitong Liu Beiping Pan Bin Xing Binghong Wu Bingxin Qu Bolin Ni Boyu Wu Chen Li Cheng Jiang Cheng Zhang Chengjun Liu Chengxu Yang Chengzhong Xu Chiyu Wang Chong Zha Daisy Yi Di Wang Fanyang Lu Fei Chen Feifei Liu Feng Zheng Guanghua Yu Guiyang Li Guohua Wang Haisheng Lin Han Liu Han Wang Hao Fei Hao Lu Haoqing Jiang Haoran Sun Haotian Zhu Huangjin Dai Huankui Chen Huawen Feng Huihui Cai Huxin Peng Jackson Lv Jiacheng Shi Jiahao Bu Jianbo Li Jianglu Hu Jiangtao Guan Jianing Xu Jianwei Cai Jiarong Zhang Jiawei Song Jie Jiang Jie Liu Jieneng Yang Jihong Zhang Jin lv Jing Zhao Jinjian Li Jinxing Liu Jun Zhao Juntao Guo Kai Wang Kan Wu Lei Fu Lei He Lei Wang Li Liu Liang Dong Liya Zhan Long Cheng Long Xu Mao Zheng Meng Liu Mengkang Hu Nanli Chen Peirui Chen Peng He Pengju Pan Pengzhi Wei Qi Yang Qi Yi Roberts Wang Rongpeng Chen Rui Sun Rui Yang Ruibin Chen Ruixu Zhou Shaofeng Zhang Sheng Zhang Shihao Xu Shuaishuai Chang Shulin Liu SiQi Wang Songjia Feng Songling Yuan Tao Zhang Tianjiao Lang Tongkai Li Wei Deng Wei Li Weichao Wang Weigang Zhang Weixuan Sun Wen Ouyang Wenxiang Jiao Wenzhi Sun Wenzhuo Jia Xiang Zhang Xiangyu He Xianshun Ren XiaoYing Zhu Xiaolong Guo Xiaoxue Li Xiaoyu Ma Xican Lu Xinhua Feng Xinting Huang Xinyu Guan Xirui Li Xu Zhang Xudong Gao Xun Luo Xuxiang Qi Yangkun Chen Yangyu Tao Yanling Xiao Yantao Mai Yanze Chen Yao Ding Yeting Yang YiFan Song Yifan Yang Yijiao Zhu Yinhe Wu Yixian Liu Yong Yang Yuanjun Cai Yuanlin Tu Yue Zhang Yufei Huang Yuhang Zhou Yuhao Jiang Yuhong Liu Yuhui Hu Yujin Lin Yun Yang Yunhao Wang Yusong Zhang Zekun Wu Zelong Zhang Zhan Yu Zhaoliang Yang Zhe Zhao Zheng Li Zhenyu Huang Zhiguang Liu Zhijiang Xu Zhiqing Kui Zhiyin Zeng Zhiyuan Xiong Zhuo Han Zifan Wu Zigang Geng Zilong Zhao Ziyan Tang Ziyuan Zhu Zonglei Zhu Zhijiang Xu
large models hybrid architecture chain-of-thought MoE long-sequence

Key Findings

Methodology

This model integrates Mamba2's linear complexity long-sequence processing with Transformer’s rich contextual understanding, employing 128 layers including Mamba2, Attention, and FFN modules. The AMF/MF block pattern balances efficiency and performance. GQA reduces KV cache overhead, and FFN uses MoE with 32 experts. Pre-trained on 16T high-quality tokens, it supports 256K context length. Post-training involves supervised fine-tuning, adaptive long-short CoT fusion, multi-round deliberation, and two-stage RL to enhance reasoning and instruction-following capabilities.

Key Results

  • Achieved 1356 score on LMSYS Chatbot Arena, ranking 7th overall, outperforming Gemini-2.0-Flash-001 (1352) and o4-mini-2025-04-16 (1345). In 23 benchmarks, averaged 77.9%, excelling in math, coding, multi-turn dialogue, and long queries, with top 5 in several categories.
  • Adaptive chain-of-thought reduces token usage by ~50%, lowering inference costs while maintaining reasoning quality. Multi-round deliberation and RL further improve complex reasoning performance.
  • Training involved multi-stage strategies, including annealing, long-context expansion, and diverse data fusion, ensuring robustness and multi-task adaptability.

Significance

This work pushes the boundaries of long-sequence processing, combining Mamba2’s efficiency with Transformer’s understanding. The adaptive reasoning mechanism allows dynamic switching between quick responses and deep analysis, addressing key bottlenecks in large-scale NLP. It offers a practical solution for deploying high-capacity models efficiently, impacting AI applications across industries such as dialogue, reasoning, and coding, while reducing costs.

Technical Contribution

Innovations include the AMF/MF block pattern, GQA attention for cache reduction, and the integration of Mamba2 with Transformer. The adaptive chain-of-thought mechanism, combined with multi-round deliberation and RL, significantly enhances reasoning depth and efficiency. The multi-stage training and long-context expansion strategies set new standards for scalable, efficient large models.

Novelty

First to deeply fuse Mamba2 with Transformer architecture, introducing a dynamic adaptive reasoning mechanism that switches between fast and deep modes based on task complexity. The AMF/MF block design and support for 256K context length are novel contributions, addressing long-sequence bottlenecks in industry applications. These innovations mark a significant step forward in scalable NLP model design.

Limitations

  • Handling extremely long documents or multi-modal data remains challenging, with increased training costs and complexity. Further optimization needed for real-time deployment.
  • The adaptive reasoning mechanism depends on teacher models and reinforcement learning, which are computationally intensive and require careful tuning.
  • Performance in low-resource languages or domain-specific tasks still needs improvement, requiring broader data collection and fine-tuning.

Future Work

Future directions include multi-modal integration, multi-task multi-language adaptation, and hardware acceleration to reduce costs. Enhancing the robustness of adaptive reasoning in extreme scenarios and expanding application domains such as robotics and embedded systems are also key goals.

AI Executive Summary

Hunyuan-TurboS exemplifies the forefront of large-scale NLP, combining Mamba2's linear complexity long-sequence processing with Transformer’s contextual richness. With 560 billion parameters across 128 layers, it employs a hybrid AMF/MF block architecture, integrating GQA attention to minimize KV cache overhead. Pre-trained on 16 trillion tokens, the model supports an unprecedented 256K context length, enabling it to understand and generate extremely long texts.

The training process involved multiple sophisticated stages. Initially, a meticulous data pipeline curated 16T high-quality tokens, followed by multi-stage pre-training with annealing and curriculum-based long-context expansion. Post-training strategies included supervised fine-tuning on 3 million instructions, adaptive long-short chain-of-thought fusion, multi-round deliberation, and a two-stage reinforcement learning process targeting reasoning and instruction-following. These innovations collectively enhanced the model’s reasoning depth, efficiency, and robustness.

In practical evaluations, Hunyuan-TurboS achieved a top-7 ranking (score 1356) on LMSYS Chatbot Arena, outperforming models like Gemini-2.0-Flash-001. Its automated benchmark scores averaged 77.9%, with outstanding performance in mathematical reasoning, coding, and multi-turn dialogues. The core technical breakthrough is its adaptive chain-of-thought mechanism, which dynamically switches between rapid and deep reasoning modes, significantly reducing token usage and inference costs.

This work addresses longstanding challenges in long-text processing and reasoning efficiency, offering a scalable, cost-effective solution for deploying large models in industry. Its architecture and training strategies set new standards, opening pathways for multi-modal, multi-task, and multi-lingual AI systems. Despite current limitations, such as handling extreme long documents and multi-modal data, ongoing research aims to further optimize performance, reduce costs, and expand application domains, heralding a new era of intelligent, efficient large models.

Deep Analysis

Background

The evolution of NLP has been driven by transformer-based models like BERT, GPT, and T5, which achieved significant breakthroughs in language understanding. However, these models face challenges in processing ultra-long sequences due to quadratic complexity, limiting their application in tasks requiring extensive context comprehension. Recent efforts introduced models like Mamba (Gu & Dao, 2023), which utilize state-space models for linear complexity, enabling longer context handling. Mixture of Experts (MoE) techniques further expanded capacity, but integrating these architectures into a scalable, efficient framework remains complex. Industry demands for models capable of understanding multi-turn dialogues, lengthy documents, and multi-modal data have accelerated research into hybrid architectures that combine the strengths of different models, aiming to balance performance, efficiency, and scalability.

Core Problem

Despite advancements, existing large models struggle with processing ultra-long texts efficiently, often incurring high computational costs and latency. The challenge lies in designing architectures that can dynamically adapt to task complexity, switching between fast responses for simple queries and deep reasoning for complex ones. Additionally, maintaining high accuracy across diverse tasks while reducing inference costs remains unresolved. Long-context handling is limited by quadratic complexity, restricting real-world applications like long-form content analysis, multi-turn dialogues, and multi-modal integration. Addressing these bottlenecks requires innovative architectures, adaptive mechanisms, and efficient training strategies to enable scalable, versatile large models.

Innovation

This work introduces several key innovations:

1) Hybrid Transformer-Mamba2 architecture with AMF/MF blocks, balancing long-sequence efficiency and contextual understanding.

2) Adaptive chain-of-thought mechanism that dynamically switches between rapid and deep reasoning modes based on task complexity, reducing token usage by ~50%.

3) Integration of GQA attention to minimize KV cache overhead, improving inference speed.

4) Support for 256K context length via curriculum-based expansion with NTK-aware positional encoding.

5) Multi-stage training pipeline combining annealing, instruction fine-tuning, multi-round deliberation, and RL, ensuring robustness and multi-task adaptability.

These innovations collectively address the core bottlenecks in long-text processing and reasoning efficiency, setting a new standard for scalable large models.

Methodology

  • �� Data pipeline: Curate 16T high-quality tokens through URL deduplication, topic classification, heuristic filtering, and semantic deduplication.
  • �� Model architecture: Design 128-layer hybrid with interleaved AMF and MF blocks, integrating Mamba2 for linear complexity, GQA attention, and MoE FFN with 32 experts.
  • �� Pre-training: Multi-phase strategy including curriculum-based long-context expansion (4K to 256K tokens), NTK-aware positional encoding, and annealing with diverse datasets.
  • �� Post-training: Conduct supervised fine-tuning on 3M instruction data, develop adaptive CoT fusion with teacher models and reinforcement learning, implement multi-round deliberation for capability refinement, and apply two-stage RL for reasoning and instruction-following.
  • �� Optimization: Use AdamW optimizer, capacity factor γ=1.5, and extensive ablation to validate design choices.

Experiments

The model was evaluated across multiple benchmarks including MMLU, HellaSwag, GSM8k, EvalPlus, and others, with various few-shot settings. It was compared against industry benchmarks like Llama-4-Maverick, DeepSeek-V3, and Qwen3. Hyperparameters included 16T tokens, AdamW optimizer, and γ=1.5. Ablation studies confirmed the effectiveness of the hybrid blocks, GQA attention, and adaptive CoT. Results demonstrated superior performance, especially in reasoning and multi-turn dialogues, validating the architecture's scalability and efficiency.

Results

Hunyuan-TurboS scored 1356 on LMSYS Arena, ranking 7th overall, outperforming models like Gemini-2.0-Flash-001. In automated benchmarks, it averaged 77.9%, excelling in math, coding, and reasoning tasks. The adaptive CoT mechanism reduced token consumption by ~50%, lowering inference costs significantly. The multi-stage training and long-context expansion contributed to its robustness and multi-task competence. These results confirm that the model effectively balances high performance with computational efficiency, addressing industry needs for scalable, versatile NLP systems.

Applications

The model is suitable for complex dialogue systems, long-form content analysis, and multi-modal applications such as video and image understanding. Its ability to process ultra-long contexts enables use in legal, scientific, and technical domains, where detailed reasoning over extensive content is required. Its efficiency makes it feasible for deployment in real-time systems, reducing operational costs and expanding accessibility across industries.

Limitations & Outlook

Handling extremely long documents or multi-modal data still presents challenges, with increased training and inference costs. The adaptive reasoning mechanism depends on teacher models and reinforcement learning, which are computationally intensive. Further work is needed to improve multi-lingual and domain-specific generalization, as well as hardware acceleration for broader deployment. These limitations highlight ongoing areas for research to fully realize the model's potential.

Plain Language Accessible to non-experts

想象你在一家大厨房里做饭,有不同的厨师:一些厨师擅长快速做简单菜肴,比如煎蛋;一些厨师则擅长复杂菜肴,比如烤火鸡。传统上,所有厨师都用同一种方式做菜,效率不高。现在,这个厨房引入了一套智能调度系统,可以根据菜的复杂程度自动安排厨师:简单的菜用快厨师,复杂的菜用慢厨师,确保既快又好。Hunyuan-TurboS就像这个智能厨房,它结合了快速厨师和深思熟虑厨师的优点,能根据问题的难度自动调整处理方式。它还能处理超长的菜单,比如长篇小说或复杂的工程图纸,不会因为内容太多而变慢。这样,这个系统既聪明又高效,能帮你节省时间,又能解决各种难题,就像一个既快又懂事的厨师团队。

ELI14 Explained like you're 14

想象你在学校里,有两个老师:一个讲得快,能快速回答简单问题;另一个讲得慢,但能帮你搞懂复杂的难题。平时你问“今天天气怎么样?”,老师会很快告诉你;但如果你问“为什么地球有四季?”,老师会花时间帮你分析。Hunyuan-TurboS就像这两个老师的结合体,它可以根据你问的问题自动选择用快的方式还是深度思考。比如你问“明天的数学考试难不难?”,它会用详细的推理帮你准备。它还能处理超长的问题,比如一篇长文章或一份复杂的工程图,不会因为内容太多而变慢。这样,它既聪明又高效,既能帮你节省时间,又能帮你解决难题,就像一个既快又懂事的超级老师。

Abstract

As Large Language Models (LLMs) rapidly advance, we introduce Hunyuan-TurboS, a novel large hybrid Transformer-Mamba Mixture of Experts (MoE) model. It synergistically combines Mamba's long-sequence processing efficiency with Transformer's superior contextual understanding. Hunyuan-TurboS features an adaptive long-short chain-of-thought (CoT) mechanism, dynamically switching between rapid responses for simple queries and deep "thinking" modes for complex problems, optimizing computational resources. Architecturally, this 56B activated (560B total) parameter model employs 128 layers (Mamba2, Attention, FFN) with an innovative AMF/MF block pattern. Faster Mamba2 ensures linear complexity, Grouped-Query Attention minimizes KV cache, and FFNs use an MoE structure. Pre-trained on 16T high-quality tokens, it supports a 256K context length and is the first industry-deployed large-scale Mamba model. Our comprehensive post-training strategy enhances capabilities via Supervised Fine-Tuning (3M instructions), a novel Adaptive Long-short CoT Fusion method, Multi-round Deliberation Learning for iterative improvement, and a two-stage Large-scale Reinforcement Learning process targeting STEM and general instruction-following. Evaluations show strong performance: overall top 7 rank on LMSYS Chatbot Arena with a score of 1356, outperforming leading models like Gemini-2.0-Flash-001 (1352) and o4-mini-2025-04-16 (1345). TurboS also achieves an average of 77.9% across 23 automated benchmarks. Hunyuan-TurboS balances high performance and efficiency, offering substantial capabilities at lower inference costs than many reasoning models, establishing a new paradigm for efficient large-scale pre-trained models.

cs.CL