Experience-Calibrated Contrastive Decoding for Mitigating Hallucinations in LM-Based Text-to-Speech

TL;DR

Experience-Calibrated Contrastive Decoding (ECCD) reduces speech hallucinations, lowering WER/CER by up to 55.6%.

eess.AS 🔴 Advanced 2026-08-01 192 views
Chenlin Liu Minghui Fang Zhonghao Bi Zekai Su Rong Wang Jiqing Han
Speech Synthesis Contrastive Decoding Hallucination Mitigation Decoding Control Multilingual

Key Findings

Methodology

This work introduces a conditional information perspective, distinguishing text-derived alignment from experience-based acoustic context and speech regularities. Using predictions from the same speech LM with and without text conditions, ECCD is designed as a training-free method that enhances alignment support while preserving valuable experience information. It applies only positive alignment reinforcement, calibrated by set-level experience compatibility (ECC). The approach is validated across four models on SeedTTS-Eval and CV3-Eval datasets, showing up to 55.6% reduction in WER/CER, with maintained speaker similarity and naturalness. The method dynamically balances alignment and experience influences during decoding, addressing the onset and propagation of hallucinations.

Key Results

  • Across four models, ECCD achieves an average WER/CER reduction of over 50%, with the highest reaching 55.6%. It significantly improves performance in multilingual and long-form scenarios, with a CMOS gain of +0.644 and strong speaker similarity. The analysis reveals that alignment influence varies at linguistic boundaries, especially at initial error points, validating the targeted enhancement strategy. Compared to baseline sampling, ECCD offers more stable and consistent improvements, especially in complex content.
  • Quantitative metrics demonstrate that ECCD outperforms traditional sampling and architecture-based methods, reducing hallucinations and content errors. Human listening tests confirm perceptual improvements, with increased naturalness and content fidelity. The method's adaptability across languages and models highlights its broad applicability, especially in low-resource or challenging scenarios.
  • Ablation studies show that the set-level ECC effectively calibrates the enhancement strength, preventing overcorrection and maintaining experience information. The approach effectively suppresses hallucinations while preserving speech fluency, demonstrating its potential for real-time, decoding-time control in practical TTS systems.

Significance

This research advances the field of neural TTS by shifting the focus from架构或训练优化到解码时的条件信息调控。通过强化对齐支持,显著改善长句和多语种内容的准确性和自然度,解决了传统方法难以应对的幻觉问题。其无训练、实时调节的特性,为未来端到端语音系统的鲁棒性提供了新思路。该技术不仅提升了语音合成的内容一致性,也为多模态、多任务场景中的内容控制奠定了基础。长远来看,它有望推动智能语音交互的普及,改善人机沟通的自然性和可信度。

Technical Contribution

本文创新性地将对比解码引入语音生成,结合集合级体验兼容性系数(ECC)实现动态调节。区别于传统的采样或架构优化,提出无训练的解码时调控机制,强调在每一步实时强化对齐支持,兼顾经验信息的保留。通过专家分布和经验代理的结合,确保生成内容既符合文本,也自然流畅。该机制在多模型、多语种环境下验证了其广泛适用性,为解码策略提供了理论和工程创新,推动了语音合成内容控制的边界。

Novelty

这是首次将条件信息区分引入LM-based TTS的解码控制,提出正向增强策略,区别于以往依赖架构或训练优化的方案。利用集合级体验兼容性系数实现自适应调节,确保在不同生成阶段平衡对齐和经验信息,显著减少幻觉。该方法实现了无训练、实时调控,为语音生成提供了全新思路,解决了长文本和多语种环境中的幻觉难题,具有开创性意义。

Limitations

  • 该方法依赖声学模型的预测能力,在极端噪声或语音变异场景下可能效果有限。调节参数需要经验调优,且在特殊内容(如诗歌、方言)中表现不足。
  • 在极端长文本或复杂语境中,仍存在一定的幻觉风险,未来需结合多模态信息进行优化。调节机制可能在某些场景下调节过度或不足。
  • 算法在极端低资源或极端语音变异条件下的鲁棒性仍需验证,未来需结合自适应机制提升泛化能力。

Future Work

未来将探索多模态条件信息融合,如视觉、上下文等,增强对齐支持的鲁棒性。结合强化学习和自适应调节机制,提升系统的动态调控能力。计划在实际工业场景中测试,推动端到端系统的应用落地。还将研究多任务、多场景的泛化能力,提升系统的适应性和稳定性。

AI Executive Summary

Deep Dive

Plain Language Accessible to non-experts

想象你在烘焙一块蛋糕,厨师需要不断调整火候和调料,确保每一层都完美。语音合成也是如此,模型在生成每个音节时,要在内容准确和自然流畅之间找到平衡。如果只关注内容,可能会偏离目标,就像蛋糕太咸;只关注自然,可能内容不对,就像蛋糕没有味道。科学家们设计了一种“智能调味料”,在模型每次拼接声音时,实时加强对内容的支持,避免偏离。这样,生成的语音既符合文本,又听起来自然,就像厨师掌握了火候,做出美味的蛋糕一样。这种方法让合成的语音更像人说话,不偏离目标,也更自然流畅,未来可以让智能助手、导航语音变得更聪明、更贴心!

ELI14 Explained like you're 14

想象你在玩一个超级复杂的拼图游戏,每次拼完一块都要确保它和整体图案对得上。语音合成也是这样,模型每次生成一个声音片段时,要确保它既符合你想要的内容,又听起来像自然的说话。可是,有时候模型会偏离主题,像拼图拼错了地方,虽然听起来还挺自然,但内容错了。科学家们发现,这种偏离其实是因为模型在平衡“内容对齐”和“自然流畅”这两件事时出了问题。于是,他们设计了一种聪明的“调味料”,在模型拼每块拼图时,实时加强对内容的支持,让偏离变少。这样,生成的语音既准确又自然,就像厨师用心调味,做出美味佳肴一样。这个新方法让语音变得更像人说话,不会偏离目标,也更自然流畅,未来可以让智能助手、导航语音变得更聪明、更贴心!

Abstract

Language model-based text-to-speech (LM-based TTS) remains vulnerable to speech hallucinations that deviate from the target text. Existing mitigation mainly relies on architectural changes or additional training, while decoding-time control remains underexplored. We present a conditional information view that distinguishes text-derived alignment information from experience information supplied by acoustic context and learned speech regularities. We hypothesize that an important class of hallucinations begins when alignment support is insufficiently reflected in the selected token at a vulnerable transition. Using predictions from the same speech LM with and without text conditions, we propose Experience-Calibrated Contrastive Decoding (ECCD), a training-free method that strengthens alignment support while preserving useful experience information. ECCD preserves the original expert distribution, applies only positive alignment enhancement, and calibrates its strength using set-level experience compatibility. Across four models, ECCD reduces WER/CER by up to 55.6% in all SeedTTS-Eval settings and 24 of 25 multilingual CV3-Eval settings. A listening test yields a CMOS gain of $+0.644$ while retaining strong speaker similarity. Further analysis shows that alignment influence and decision-level gain vary within linguistic units and are lower at first-error boundaries than at matched correct boundaries. Overall, these extensive experiments and analyses identify conditional information control as a promising decoding-time direction for mitigating speech hallucination.

eess.AS cs.LG eess.SP