SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment

TL;DR

SongBench introduces a 7-dimension framework with 11,717 samples, achieving superior song quality evaluation precision.

eess.AS 🔴 Advanced 2026-04-16 28 views
Dapeng Wu Shun Lei Wei Tan Guangzheng Li Yunzhe Wang Huaicheng Zhang Lishi Zuo Zhiyong Wu
text-to-song music evaluation multi-dimensional analysis deep learning dataset

Key Findings

Methodology

SongBench defines seven core dimensions (Vocal, Instrument, Melody, Structure, Arrangement, Mixing, Musicality) and uses an expert-annotated dataset to reduce inter-dimensional interference and ensure reliable evaluation.

Key Results

  • Experiments show SongBench achieves an LCC of 0.976 in Musicality, outperforming SongEval's 0.839.
  • At the system level, SongBench captures subtle improvements in commercial models, e.g., Suno v4.5 to v5 scores rising from 6.60 to 6.86.
  • AB tests demonstrate SongBench achieves over 60% accuracy in intra-model comparisons, far exceeding SongEval's 43.48%.

Significance

SongBench addresses score compression and dimensional overlap issues in existing benchmarks, providing a refined tool for evaluating AI-generated music quality. It has significant implications for academia and industry, guiding professional music generation systems.

Technical Contribution

SongBench introduces a novel evaluation framework grounded in core musical elements, combined with expert annotations, significantly improving automated evaluation reliability and resolution.

Novelty

SongBench is the first fine-grained evaluation framework dedicated to text-to-song generation, reducing dimensional overlap and achieving higher precision through expert annotations.

Limitations

  • The framework relies on expert annotations, which may incur high costs for new domains.
  • Its evaluation capability for non-mainstream music styles remains unverified.

Future Work

Future work could explore automated annotation techniques to reduce dependency on experts and expand the framework to cover more music styles and languages.

AI Executive Summary

Recent advancements in text-to-song generation have enabled realistic musical content production, yet existing benchmarks fail to capture multi-dimensional aesthetic nuances. SongBench introduces a seven-dimension evaluation framework, covering Vocal, Instrument, Melody, Structure, Arrangement, Mixing, and Musicality, paired with 11,717 expert-annotated samples to achieve superior reliability and resolution.

Experimental results show SongBench aligns closely with expert ratings, achieving an LCC of 0.976 in Musicality, significantly outperforming SongEval. The framework effectively distinguishes subtle performance differences, such as Suno v4.5 to v5 score improvements.

SongBench provides a robust tool for academia and industry to evaluate and refine AI-generated music systems. Future research could focus on automating annotations and expanding the framework to diverse music styles and languages.

Deep Analysis

Background

Text-to-song generation has advanced rapidly, but benchmarks like SongEval and MusicEval suffer from dimensional overlap and score compression, limiting their ability to evaluate high-quality models.

Core Problem

Existing benchmarks struggle to distinguish subtle performance differences, especially among high-quality models, due to score saturation and insufficient resolution.

Innovation

SongBench introduces a seven-dimension evaluation framework based on core musical elements, reducing inter-dimensional interference and improving evaluation precision through expert annotations.

Methodology

  • �� Define seven dimensions: Vocal, Instrument, Melody, Structure, Arrangement, Mixing, Musicality.
  • �� Construct an expert-annotated dataset of 11,717 samples from diverse models and real music.
  • �� Employ double-blind annotation and rigorous quality control to ensure consistency.
  • �� Use MuQ as a self-supervised framework for model training.

Experiments

Experiments include ID and OOD test sets to evaluate model performance at system and sample levels. Metrics include LCC, SRCC, and KTAU to measure alignment with expert ratings.

Results

Results show SongBench achieves an LCC of 0.976 in Musicality, outperforming SongEval. It effectively distinguishes subtle improvements, such as Suno v4.5 to v5 score changes.

Applications

SongBench can be used to evaluate and optimize music generation models, helping developers identify bottlenecks and guide improvements.

Limitations & Outlook

The framework relies on expert annotations, which may incur high costs. Its ability to evaluate non-mainstream music styles remains untested.

Plain Language Accessible to non-experts

Imagine a factory producing meals. Current music generation models are like automated cooking machines that quickly prepare food but lack finesse. SongBench acts like a team of expert food critics, scoring meals on ingredients, cooking techniques, and presentation, helping the chef improve the dishes.

ELI14 Explained like you're 14

Think of a music game where you create songs for different levels. SongBench is like the scoring system in the game, rating your work on melody, vocals, and arrangement, showing you where to improve for a perfect score!

Glossary

LCC (Linear Correlation Coefficient)

Measures the linear correlation between model predictions and expert ratings.

Used to evaluate SongBench's prediction accuracy.

Musicality

Evaluates the overall artistic impact and auditory pleasure of a song.

A key dimension in the framework.

MuQ (Music Representation Learning)

A self-supervised framework for extracting music features.

Used to train SongBench's evaluation model.

Score Compression

Scores clustering in high ranges, reducing differentiation between models.

A major issue in existing benchmarks.

Double-Blind Annotation

Samples are presented without revealing their origin to reduce bias.

Ensures objective annotation results.

Open Questions Unanswered questions from this research

  • 1 How can expert dependency be reduced for cost efficiency?
  • 2 Can the framework evaluate non-mainstream music styles effectively?

Applications

Immediate Applications

Model Optimization

Helps developers identify performance bottlenecks and guide improvements.

Music Quality Evaluation

Provides a professional tool for assessing AI-generated music quality.

Long-term Vision

Automated Annotation

Develop techniques to reduce reliance on expert annotations.

Abstract

Recent advancements in Text-to-Song generation have enabled realistic musical content production, yet existing evaluation benchmarks lack the professional granularity to capture multi-dimensional aesthetic nuances. In this paper, we propose SongBench, a specialized framework for fine-grained song assessment across seven key dimensions: Vocal, Instrument, Melody, Structure, Arrangement, Mixing, and Musicality. Utilizing this framework, we construct an expert-annotated database comprising 11,717 samples from state-of-the-art models, labeled by music professionals. Extensive experimental results demonstrate that SongBench achieves high correlation with expert ratings. By revealing fine-grained performance gaps in current state-of-the-art models, SongBench serves as a diagnostic benchmark to steer the development toward more professional and musically coherent song generation.

eess.AS cs.AI cs.SD