GDPO-Listener: Expressive Interactive Head Generation via Auto-Regressive Flow Matching and Group reward-Decoupled Policy Optimization

TL;DR

GDPO-Listener generates expressive head motions via Auto-Regressive Flow Matching and Group reward-Decoupled Policy Optimization.

cs.CV 🔴 Advanced 2026-03-26 6 views
Zhangyu Jin Maksim Siniukov Deuksin Kwon Ashutosh Chaubey Mohammad Soleymani
3D head generation Auto-Regressive Flow Policy Optimization FLAME parameters Semantic control

Key Findings

Methodology

The paper introduces the GDPO-Listener framework, which uses Auto-Regressive Flow Matching (AR-Flow) and Group reward-Decoupled Policy Optimization (GDPO) to generate expressive head motions. AR-Flow enables stable supervised learning, while GDPO incentivizes high-variance expressive generations by isolating reward normalization across distinct FLAME parameter groups. The framework also supports explicit semantic text control.

Key Results

  • On the Seamless Interaction dataset, GDPO-Listener outperforms existing baselines in long-term kinematic variance, visual expressivity, and semantic controllability.
  • On the DualTalk dataset, GDPO-Listener's listener dynamic deviation (FDD) is significantly lower than the DualTalk baseline, indicating its advantage in generating dynamic reactions.
  • Ablation studies confirm GDPO's critical role in enhancing expressivity and dynamic consistency.

Significance

GDPO-Listener is significant in virtual human synthesis, addressing the 'Regression-to-the-Mean' problem in listener motion generation, enhancing motion naturalness and expressivity. The framework maintains stable dynamics during long-sequence inference and supports semantic text control, providing greater expressive freedom for virtual human interaction.

Technical Contribution

GDPO-Listener overcomes limitations of existing methods by introducing AR-Flow and GDPO, offering new theoretical guarantees and engineering possibilities. The expanded FLAME parameter space supports more complex nonverbal motion generation, significantly enhancing the model's expressive capabilities.

Novelty

GDPO-Listener is the first to introduce Group reward-Decoupled Policy Optimization in listener head generation, overcoming the limitations of traditional supervised learning and addressing expressivity issues in one-to-many generative tasks.

Limitations

  • In some complex contexts, the model may still generate unnatural motions, especially when training data is imbalanced.
  • Dependence on FLAME parameters may limit the model's applicability to other 3D models.

Future Work

Future research could explore applications on more 3D models, further optimize GDPO strategies to enhance generation quality, and validate its generality in more contexts.

AI Executive Summary

Generating realistic 3D head motion for virtual human synthesis is a significant challenge. Existing methods achieve impressive results for speaking heads but often face the 'Regression-to-the-Mean' problem in listener motions, resulting in static faces and lacking complex nonverbal motion space. The GDPO-Listener framework addresses these issues through Auto-Regressive Flow Matching and Group reward-Decoupled Policy Optimization. Auto-Regressive Flow Matching provides stable supervised learning, while Group reward-Decoupled Policy Optimization incentivizes high-variance expressive generations by isolating reward normalization across distinct FLAME parameter groups. Additionally, the framework supports explicit semantic text control to ensure contextually appropriate responses. Extensive evaluations demonstrate that GDPO-Listener outperforms existing baselines in long-term kinematic variance, visual expressivity, and semantic controllability. However, the model may still generate unnatural motions in some complex contexts. Future research could explore applications on more 3D models and further optimize GDPO strategies to enhance generation quality.

Deep Analysis

Background

The field of virtual human synthesis has seen significant advancements, particularly in speaking head generation. However, listener motion generation remains challenging, with existing methods often facing the 'Regression-to-the-Mean' problem, resulting in static faces and lacking complex nonverbal motion space.

Core Problem

Listener head generation is a non-deterministic, one-to-many problem. Existing methods often treat it as a standard supervised learning task, leading models to generate low-variance 'mean' motions.

Innovation

GDPO-Listener addresses expressivity issues in listener head generation through Auto-Regressive Flow Matching and Group reward-Decoupled Policy Optimization. Auto-Regressive Flow Matching provides stable supervised learning, while Group reward-Decoupled Policy Optimization incentivizes high-variance expressive generations by isolating reward normalization across distinct FLAME parameter groups.

Methodology

  • �� Auto-Regressive Flow Matching (AR-Flow): for stable supervised learning.
  • �� Group reward-Decoupled Policy Optimization (GDPO): incentivizes high-variance expressive generations by isolating reward normalization.
  • �� Expanded FLAME parameter space: supports more complex nonverbal motion generation.
  • �� Explicit semantic text control: ensures contextually appropriate responses.

Experiments

Extensive evaluations on the Seamless Interaction and DualTalk datasets demonstrate GDPO-Listener's superiority in long-term kinematic variance, visual expressivity, and semantic controllability.

Results

GDPO-Listener's listener dynamic deviation (FDD) on the Seamless Interaction dataset is significantly lower than the DualTalk baseline, indicating its advantage in generating dynamic reactions.

Applications

GDPO-Listener can be applied in virtual human interaction, gaming animation, and film production, providing greater expressive freedom and naturalness.

Limitations & Outlook

The model may still generate unnatural motions in some complex contexts, especially when training data is imbalanced. Future research could explore applications on more 3D models.

Plain Language Accessible to non-experts

Imagine a stage play where actors need to make natural facial expressions and movements based on dialogue. GDPO-Listener acts like a smart director, helping actors make appropriate reactions when hearing dialogue. By analyzing the semantics and emotions of the dialogue, GDPO-Listener can generate a rich variety of facial motions, just like actors making different performances based on the plot.

ELI14 Explained like you're 14

Imagine you're playing a game where characters need to make expressions based on dialogue. GDPO-Listener is like an AI assistant in the game, helping characters make appropriate reactions when hearing dialogue. It analyzes the emotions and semantics of the dialogue to generate a rich variety of facial motions, just like you make different choices in the game based on the plot.

Glossary

Auto-Regressive Flow Matching

A framework for stable supervised learning, generating intermediate states through linear interpolation.

Used for generating high-fidelity base policy.

Group reward-Decoupled Policy Optimization

Incentivizes high-variance expressive generations by isolating reward normalization across distinct parameter groups.

Used to optimize expressivity in listener head generation.

FLAME parameters

A 3D morphable model parameterizing facial motion for nonverbal actions.

Expanded parameter space supports complex motion generation.

Semantic text control

Controls generated facial motions through explicit semantic text input.

Ensures contextually appropriate responses.

Dynamic Deviation

Evaluates the deviation in dynamic consistency between generated and real motions.

Used to assess the naturalness of listener head generation.

Open Questions Unanswered questions from this research

  • 1 How can GDPO-Listener be applied to more 3D models? Current methods' reliance on FLAME parameters may limit its applicability.
  • 2 How can GDPO strategies be further optimized to enhance generation quality? Current methods may still generate unnatural motions in some complex contexts.

Applications

Immediate Applications

Virtual Human Interaction

GDPO-Listener can enhance the naturalness and expressivity of virtual assistants, providing a better user experience.

Gaming Animation

Can be used to generate more natural game character motions, enhancing game immersion.

Long-term Vision

Film Production

GDPO-Listener can generate realistic character animations, reducing the cost and time of manual animation production.

Abstract

Generating realistic 3D head motion for dyadic interactions is a significant challenge in virtual human synthesis. While recent methods achieve impressive results with speaking heads, they frequently suffer from the `Regression-to-the-Mean' problem in listener motions, collapsing into static faces, and lack the parameter space for complex nonverbal motions. In this paper, we propose GDPO-Listener, a novel framework that achieves highly expressive speaking and listening motion generation. First, we introduce an Auto-Regressive Flow Matching architecture enabling stable supervised learning. Second, to overcome kinematic stillness, we apply the Group reward-Decoupled Policy Optimization (GDPO). By isolating reward normalization across distinct FLAME parameter groups, GDPO explicitly incentivizes high variance expressive generations. Finally, we enable explicit semantic text control for customizable responses. Extensive evaluations across the Seamless Interaction and DualTalk datasets demonstrate superior performance compared to existing baselines on long-term kinematic variance, visual expressivity and semantic controllability.

cs.CV