AVENUE: Audio-Video EditiNg Understanding and Evaluation

TL;DR

AVENUE evaluates audio-video editing models' modality-selectivity using 1,291 clips and 7,957 instructions.

cs.MM 🔴 Advanced 2026-08-31 8 views
Hayeon Kim Yoojin Jang Jaejun Yoo
audio-video editing multimodal dataset model evaluation modality-selectivity

Key Findings

Methodology

AVENUE provides a benchmark with 1,291 source clips and 7,957 editing instructions covering audio, video, and AV-coupled edits. The evaluation framework specifies intended changes and content to remain intact for each sample, supporting modality-selectivity assessment.

Key Results

  • Existing models often induce unintended changes in the non-target modality during editing.
  • Joint models excel in AV-coupled editing in terms of accuracy and consistency.
  • Separate models perform better in isolated audio and video editing.

Significance

This study fills a gap in evaluating audio-video editing models by providing a comprehensive benchmark and evaluation framework, driving progress toward more controllable AV editing models.

Technical Contribution

AVENUE provides the first systematic analysis of modality-selectivity across editing paradigms and introduces the Selective Controllability metric to assess models' ability to preserve unintended modalities.

Novelty

AVENUE is the first benchmark to offer comprehensive edit-type and modality coverage, addressing the modality-blind and sample-agnostic issues of existing evaluation systems.

Limitations

  • Existing models struggle to avoid unintended impacts on the non-target modality during editing.
  • The evaluation framework's applicability across different datasets needs further validation.

Future Work

Future work could explore finer modality-selectivity control mechanisms and extend the benchmark to more diverse datasets.

AI Executive Summary

Audio-video editing is a complex task requiring models to modify audio and video content according to a target prompt. Existing evaluation systems often overlook modality-selectivity and sample specificity, making it difficult to effectively assess models' editing capabilities. AVENUE addresses this gap by providing a benchmark with 1,291 source clips and 7,957 editing instructions. This benchmark covers audio, video, and AV-coupled edit types and offers a sample-specific, modality-aware evaluation framework that specifies both the intended changes and the content that must remain intact for each sample.

Through evaluating representative audio-video editing models, the study finds that existing models frequently induce unintended changes in the non-target modality during editing. AVENUE's evaluation framework introduces the Selective Controllability metric, providing the first systematic analysis of modality-selectivity across different editing paradigms.

The significance of AVENUE lies in its potential to drive the development of more controllable audio-video editing models, offering a new standard for evaluation in both academia and industry. Future work could explore finer modality-selectivity control mechanisms and extend the benchmark to more diverse datasets.

Deep Analysis

Background

The research background covers recent advances in multimodal generative models, which extend beyond content creation to content editing tasks. Existing audio-video editing frameworks propose diverse architectures but face challenges in inferring modality-selective edit scopes.

Core Problem

Audio-video editing requires models to determine not only what should change but also which modality should be preserved. Existing evaluation systems often overlook modality-selectivity and sample specificity, making it difficult to effectively assess models' editing capabilities.

Innovation

AVENUE's core innovation lies in providing a comprehensive benchmark and evaluation framework that covers diverse edit types and modality combinations, addressing the modality-blind and sample-agnostic issues of existing evaluation systems.

Methodology

  • �� Provides 1,291 source clips and 7,957 editing instructions
  • �� Covers audio, video, and AV-coupled edit types
  • �� Sample-specific, modality-aware evaluation framework
  • �� Introduces Selective Controllability metric

Experiments

The experimental design includes evaluating representative audio-video editing models across joint, sequential, and separate paradigms. Evaluation metrics include Edit Accuracy, Modality Selectivity, Perceptual Quality, and AV Consistency.

Results

The study finds that existing models frequently induce unintended changes in the non-target modality during editing. Joint models excel in AV-coupled editing, while separate models perform better in isolated editing.

Applications

AVENUE's application scenarios include driving the development of more controllable audio-video editing models, offering a new standard for evaluation in both academia and industry.

Limitations & Outlook

Existing models struggle to avoid unintended impacts on the non-target modality during editing. The evaluation framework's applicability across different datasets needs further validation.

Plain Language Accessible to non-experts

Imagine you're in a kitchen. Audio is like the soup you're cooking, and video is like the bread you're baking. You want to add some salt to the soup but don't want the bread's flavor to change. Audio-video editing is like this process, where the model needs to know when to change only the soup's flavor without affecting the bread. AVENUE is like a recipe that tells you how to perfectly add salt to the soup without affecting the bread.

ELI14 Explained like you're 14

Imagine you're playing a game where you can change a character's voice or appearance. You want to make the voice cooler but don't want the appearance to change. Audio-video editing is like this, where the model needs to know when to change only the voice without affecting the appearance. AVENUE is like a guide that helps you perfectly adjust the voice without changing the appearance.

Glossary

Audio-Video Editing

The process of modifying audio and video content according to a target prompt.

Refers to the editing task performed by models based on prompts in the paper.

Modality Selectivity

The ability of a model to selectively change one modality while preserving another during editing.

Evaluates whether a model can maintain the non-target modality unchanged during editing.

Selective Controllability

Assesses a model's ability to preserve the integrity of the unintended modality.

Used as an evaluation metric to measure editing precision.

Joint Model

A unified framework that processes audio and visual streams simultaneously.

Refers to the AvED model in the paper.

Separate Model

Single-modality models applied independently to audio and video.

Refers to the RAVE+ZETA and RAVE+SDEdit models in the paper.

Open Questions Unanswered questions from this research

  • 1 How to completely avoid unintended impacts on the non-target modality during editing?
  • 2 How applicable is the evaluation framework across different datasets?

Applications

Immediate Applications

Video Editing Software

Can be used to develop more refined audio-video editing tools, helping users achieve more precise editing effects.

Long-term Vision

Intelligent Media Production

Advances the development of intelligent media production, enabling more efficient content creation and editing.

Abstract

Audio-video (AV) editing aims to modify audio and video content according to a target prompt. Unlike single-modality editing, AV editing requires models to infer a modality-selective edit scope from the prompt alone: determining not only what should change, but also which modality should be preserved. Faithfully evaluating such models therefore requires both (i) benchmarks that span diverse edit types and modality categories, and (ii) evaluation that is itself modality-aware and sample-specific. However, existing AV editing benchmarks provide limited coverage of edit types and modality combinations, while current evaluation systems are often modality-blind and sample-agnostic, making it difficult to assess whether models faithfully preserve the unintended modality. To address these gaps, we introduce AVENUE, Audio-Video EditiNg Understanding and Evaluation, comprising two contributions: (1) a benchmark of 1,291 source clips and 7,957 editing instructions across audio-targeted, video-targeted, and AV-coupled edit types, curated and human-verified from VGGSound; and (2) a sample-specific, modality-aware evaluation framework that specifies, for each sample, both the intended change and the content that must remain intact. We evaluate representative AV editing models spanning three editing paradigms : joint, sequential, and separate, providing the first systematic analysis of modality-selectivity across paradigms. Our findings reveal a fundamental open challenge: when editing one modality, existing models frequently induce unintended changes in the other, regardless of paradigm. AVENUE provides a benchmark and modality-aware evaluation framework to drive progress toward more controllable AV editing models. Our dataset is publicly available on Hugging Face: https://huggingface.co/datasets/AVENUE-dataset/AVENUE.

cs.MM cs.CV cs.SD