Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning
The study reveals quality issues in multilingual speech datasets, emphasizing the need for sociolinguistic awareness and proactive language planning.
Key Findings
Methodology
The study employs both quantitative and qualitative methods to assess the quality of Mozilla Common Voice 17.0, FLEURS, and Vox Populi datasets. Metrics like Signal-to-Noise Ratio and Voice Activity Detection are used, along with native speaker reviews.
Key Results
- In MCV17's nan_tw subset, serious quality issues were found, with 99% of utterances below 7 seconds, rendering the data nearly unusable.
- A positive correlation between a language's institutionalization status and its dataset quality was identified, with more severe issues in low-resource languages.
- Guidelines were proposed to improve future dataset development, emphasizing sociolinguistic awareness and language planning principles.
Significance
The study highlights the importance of data quality for downstream applications and research, especially in low-resource languages. The proposed guidelines help improve dataset quality and facilitate community-led language planning and revitalization.
Technical Contribution
Provides a comprehensive analysis of quality issues in multilingual speech datasets and offers improvement suggestions. Emphasizes the importance of sociolinguistic factors in dataset design.
Novelty
First to systematically incorporate sociolinguistic factors into multilingual speech dataset quality assessment, proposing new language planning methods.
Limitations
- The study focuses on a limited number of datasets, which may not represent all multilingual datasets comprehensively.
- Analysis of low-resource languages requires more linguistic expert involvement.
- Does not cover all possible sociolinguistic factors.
Future Work
Future research should explore how the dataset creation process can be leveraged as a tool for community-led language planning and revitalization.
AI Executive Summary
The study reveals quality issues in multilingual speech datasets, particularly in low-resource languages. Existing datasets like Mozilla Common Voice 17.0, FLEURS, and Vox Populi suffer from significant micro and macro-level quality issues, affecting downstream application evaluation results.
The research employs both quantitative and qualitative methods, analyzing metrics like Signal-to-Noise Ratio and Voice Activity Detection, and inviting native speakers to review samples. Results show a positive correlation between a language's institutionalization status and its dataset quality, with more severe issues in low-resource languages.
Guidelines were proposed to improve future dataset development, emphasizing the importance of sociolinguistic awareness and language planning principles. The study suggests leveraging the dataset creation process as a tool for community-led language planning and revitalization.
Deep Dive
Abstract
Our quality audit for three widely used public multilingual speech datasets - Mozilla Common Voice 17.0, FLEURS, and Vox Populi - shows that in some languages, these datasets suffer from significant quality issues, which may obfuscate downstream evaluation results while creating an illusion of success. We divide these quality issues into two categories: micro-level and macro-level. We find that macro-level issues are more prevalent in less institutionalized, often under-resourced languages. We provide a case analysis of Taiwanese Southern Min (nan_tw) that highlights the need for proactive language planning (e.g. orthography prescriptions, dialect boundary definition) and enhanced data quality control in the dataset creation process. We conclude by proposing guidelines and recommendations to mitigate these issues in future dataset development, emphasizing the importance of sociolinguistic awareness and language planning principles. Furthermore, we encourage research into how this creation process itself can be leveraged as a tool for community-led language planning and revitalization.