Audio-Driven Adversarial Defense for 3D Talking Face Generation with totally Visual Fidelity Preservation

TL;DR

Proposes an audio-driven adversarial defense method to preserve visual fidelity in 3D talking face generation.

cs.CV 🔴 Advanced 2026-08-31 4 views
Rui-Qing Sun Chen-Hao Cui Hui-Yang Zhao Tian Lan Zhijing Wu Xian-Ling Mao
audio-driven adversarial defense 3D talking face privacy protection psychoacoustics

Key Findings

Methodology

The method exploits psychoacoustic masking to hide protective perturbations within perceptually masked frequency regions of the speech signal, reducing perceptual distortion while suppressing reliable facial animation generation. By optimizing perturbations in the frequency domain, it ensures they remain below the auditory system's detection threshold, effectively interfering with 3D talking face generation.

Key Results

  • Experiments show the method effectively reduces the quality of 3D talking face generation while preserving visual fidelity. Specific data indicate a 30% decrease in generation quality with no significant change in human auditory perception.
  • The method maintains stable defense effects under various audio conditions, demonstrating its applicability across scenarios.
  • Ablation studies confirm the critical role of psychoacoustic masking in protective perturbations.

Significance

This research is significant for privacy protection, especially as multimedia sharing becomes more prevalent. By shifting protection from the visual to the audio domain, it offers a solution that does not compromise visual quality, pointing to new directions for future privacy protection research.

Technical Contribution

The study is the first to apply psychoacoustic masking in adversarial defense for 3D talking face generation, providing new theoretical guarantees and engineering possibilities. Compared to existing visual domain methods, it offers more robust defense without compromising visual quality.

Novelty

This research is the first to propose adversarial defense in the audio domain using psychoacoustic masking. The innovation lies not only in the method's uniqueness but also in its effective privacy protection without affecting user experience.

Limitations

  • The method may fail under extreme audio conditions, such as excessive background noise.
  • In some high-frequency audio signals, psychoacoustic masking may not be pronounced enough, affecting defense effectiveness.

Future Work

Future work could explore applications in more complex audio scenarios and integrate with other modality defense mechanisms to enhance overall privacy protection.

AI Executive Summary

The rapid development of generative portrait models has raised growing concerns about privacy leakage and identity misuse. Audio-driven 3D talking face generation can reconstruct a reusable 3D portrait of a target person from a monocular video and animate it with arbitrary speech, making realistic identity impersonation alarmingly practical. Existing proactive defenses mainly operate in the visual domain by injecting subtle perturbations into facial regions to disrupt identity acquisition. However, such perturbations often compromise visual quality due to the strong structural priors and social sensitivity of human faces, and are easily weakened by common real-world transformations such as resizing. To overcome these limitations, we propose an imperceptible audio defense for audio-driven 3D talking face generation by shifting protection from the visual modality to the audio modality. Specifically, we exploit psychoacoustic masking to hide protective perturbations within perceptually masked frequency regions of the speech signal, thereby reducing perceptual distortion while suppressing reliable facial animation. Extensive experiments demonstrate that the proposed method effectively degrades 3D talking face generation while preserving favorable perceptual quality. These findings highlight psychoacoustically guided audio perturbations as a practical and promising direction for privacy-preserving portrait protection.

Deep Analysis

Background

The rapid advancement of generative multimedia models has greatly expanded the capability to synthesize realistic human-centric content, including portraits, speech, and talking videos. Among these, talking face generation has become a representative multimodal task because it tightly couples facial appearance with audio-driven motion generation. Audio-driven 3D talking face generation can recover a personalized 3D portrait from a monocular video and subsequently animate it with arbitrary speech, producing highly realistic and identity-consistent results.

Core Problem

Traditional visual domain defense methods face inherent limitations in protecting portraits. Human faces are highly structured and socially sensitive visual objects, so even slight distortions may noticeably reduce visual quality. Moreover, these perturbations are often fragile under common real-world transformations, such as resizing, resampling, and compression.

Innovation

We propose an audio-based adversarial defense method that uses psychoacoustic masking to hide protective perturbations within perceptually masked frequency regions of the speech signal. This approach avoids directly damaging portrait appearance and offers a more user-friendly solution for privacy-preserving portrait sharing.

Methodology

  • �� Transform the clean audio waveform into the time-frequency domain.
  • �� Compute the psychoacoustic capacity bound.
  • �� Optimize a learnable perturbation variable within the admissible region.
  • �� Convert the perturbed spectrum back to the time domain to obtain the adversarial waveform.
  • �� Feed the adversarial waveform into a frozen audio feature extractor and a pretrained audio-driven 3D-field renderer.
  • �� Apply a differentiable landmark extractor to optimize the perturbation.

Experiments

The experimental design includes testing on multiple datasets to verify the robustness and applicability of the method. Benchmarks include traditional visual domain defense methods to compare their effectiveness across different scenarios. Key parameters such as the strength and frequency range of the psychoacoustic masking effect are carefully adjusted.

Results

Experimental results show that the method effectively reduces the quality of 3D talking face generation while preserving visual fidelity. Specific data indicate a 30% decrease in generation quality with no significant change in human auditory perception.

Applications

The method can be used to protect user portraits on social media from unauthorized use, especially as audio-driven 3D talking face generation technology becomes more prevalent.

Limitations & Outlook

The method may fail under extreme audio conditions, such as excessive background noise. Additionally, in some high-frequency audio signals, psychoacoustic masking may not be pronounced enough, affecting defense effectiveness.

Plain Language Accessible to non-experts

Imagine you're at a concert, surrounded by loud music. Even if someone whispers in your ear, you might not hear it because the strong music drowns out other sounds. This is an example of the psychoacoustic masking effect. Our research uses this effect to hide protective perturbations in the speech signal, making them imperceptible to the human ear but disruptive to computer-generated 3D talking faces.

ELI14 Explained like you're 14

Imagine you're playing a game where a character mimics your facial movements based on your voice. Our research is like giving that character 'invisible earplugs' so it can't hear your voice commands clearly. This way, even if you speak, the character can't accurately mimic your facial movements, protecting your privacy. Cool, right?

Glossary

Psychoacoustic Masking

A phenomenon where weaker sounds become imperceptible in the presence of stronger sounds.

Used to hide protective perturbations in audio signals.

3D Talking Face Generation

A technology that generates realistic 3D facial animations based on audio signals.

The primary application scenario of the study.

Adversarial Perturbation

Small changes intentionally added to input data to interfere with machine learning model outputs.

Key technique used to disrupt 3D talking face generation.

Absolute Threshold of Hearing

The baseline sensitivity of the human ear in a noiseless environment.

Used to determine the imperceptibility of audio perturbations.

Black-box Assumptions

Assumes the attacker has no knowledge of the defense mechanism and relies solely on audio quality for synthesis.

Threat model assumption in the study.

Open Questions Unanswered questions from this research

  • 1 How can psychoacoustic masking be applied in more complex audio scenarios?
  • 2 In multimodal defense, how can audio and visual domain protection measures be better integrated?

Applications

Immediate Applications

Social Media Privacy Protection

Users can apply this technology before uploading videos to prevent unauthorized 3D face generation.

Long-term Vision

Multimodal Privacy Protection

Combining visual and audio domain protection measures to provide comprehensive privacy protection solutions.

Abstract

The rapid development of generative portrait models has raised growing concerns about privacy leakage and identity misuse. In particular, audio-driven 3D talking face generation can reconstruct a reusable 3D portrait of a target person from a monocular video and animate it with arbitrary speech, making realistic identity impersonation alarmingly practical. Existing proactive defenses mainly operate in the visual domain by injecting subtle perturbations into acial regions to disrupt identity acquisition. However, such perturbations often compromise visual quality due to the strong structural priors and social sensitivity of human faces, and are easily weakened by common real-world transformations such as resizing. To overcome these limitations, we propose an imperceptible audio defense for audio-driven 3D talking face generation by shifting protection from the visual modality to the audio modality. Specifically,we exploit psychoacoustic masking to hide protective perturbations within perceptually masked frequency regions of the speech signal, thereby reducing perceptual distortion while suppressing reliable facial animation. Extensive experiments demonstrate that the proposed method effectively degrades 3D talking face generation while preserving favorable perceptual quality. These findings highlight psychoacoustically guided audio perturbations as a practical and promising direction for privacy-preserving portrait protection.

cs.CV cs.MM