Ramsa: A Large Sociolinguistically Rich Emirati Arabic Speech Corpus for ASR and TTS
Ramsa corpus provides baseline ASR and TTS for Emirati Arabic; Whisper-large-v3-turbo excels.
Key Findings
Methodology
The Ramsa corpus consists of 41 hours of Emirati Arabic speech from structured interviews and TV shows. A 10% subset was used to evaluate ASR and TTS models in a zero-shot setting, employing Whisper-large-v3-turbo and MMS-TTS-Ara for baseline testing.
Key Results
- Whisper-large-v3-turbo achieved the best ASR performance with a word error rate of 0.268 and a character error rate of 0.144.
- MMS-TTS-Ara excelled in TTS with a word error rate of 0.285 and a character error rate of 0.081.
- Baseline results show competitive performance but indicate room for improvement in Emirati Arabic.
Significance
The Ramsa corpus provides rich sociolinguistic data for Emirati Arabic, addressing gaps in low-resource language technologies. It supports ASR and TTS research, particularly in dialectal and gender diversity.
Technical Contribution
Ramsa surpasses existing resources in gender and dialect diversity, offering new baselines for ASR and TTS. It demonstrates model performance in zero-shot settings, providing references for future improvements.
Novelty
Ramsa is the first corpus to focus on diverse Emirati Arabic dialects and gender representation, offering broader sociolinguistic coverage than existing corpora.
Limitations
- The corpus size is limited to 41 hours, affecting model generalization.
- Baseline tests were conducted on only 10% of the dataset, which may not be comprehensive.
Future Work
Future work includes expanding the corpus size, increasing dialect and gender representation, and conducting comprehensive model evaluations on larger datasets.
AI Executive Summary
The Ramsa corpus is a developing Emirati Arabic speech corpus designed to support sociolinguistic research and low-resource language technologies. It contains 41 hours of speech data from 157 speakers, covering urban, Bedouin, and Mountain/Shihhi dialects. A 10% subset was used to evaluate ASR and TTS models, with Whisper-large-v3-turbo excelling in ASR and MMS-TTS-Ara in TTS. While baseline results are competitive, there is room for improvement.
The development of the Ramsa corpus fills a gap in low-resource language technologies for Emirati Arabic. It provides rich sociolinguistic data for ASR and TTS research, particularly in dialectal and gender diversity. Results indicate that existing models still have room for improvement in Emirati Arabic.
Future research directions include expanding the corpus size, increasing dialect and gender representation, and conducting comprehensive model evaluations on larger datasets. This will help improve ASR and TTS model performance in Emirati Arabic and advance low-resource language technologies.
Deep Analysis
Background
Arabic dialect corpora remain limited in scale and sociolinguistic detail, especially for Gulf dialects. Existing corpora like ADI17 and Alsanaa are limited in size and representativeness. The Ramsa corpus aims to fill these gaps by providing broader sociolinguistic representation.
Core Problem
Emirati Arabic is underrepresented in speech technologies, with existing corpora limited in size and diversity. This limits ASR and TTS model performance and generalization capabilities.
Innovation
The Ramsa corpus innovates by offering diverse dialect and gender representation. It covers urban, Bedouin, and Mountain/Shihhi dialects and increases female participation, providing comprehensive sociolinguistic data.
Methodology
- �� Collect 41 hours of Emirati Arabic speech data
- �� Cover 157 speakers across multiple dialects
- �� Use 10% of the dataset for ASR and TTS baseline evaluation
- �� Employ Whisper-large-v3-turbo and MMS-TTS-Ara models for testing
Experiments
Experiments used 10% of the Ramsa corpus for ASR and TTS baseline evaluation. Whisper-large-v3-turbo and MMS-TTS-Ara models were tested, excelling in ASR and TTS tasks, respectively.
Results
Whisper-large-v3-turbo achieved a word error rate of 0.268 and a character error rate of 0.144 in ASR. MMS-TTS-Ara achieved a word error rate of 0.285 and a character error rate of 0.081 in TTS, showing competitive performance in Emirati Arabic.
Applications
The Ramsa corpus can be used for ASR and TTS research, particularly in low-resource language technologies. It provides baselines for model performance across diverse dialects and genders.
Limitations & Outlook
The corpus size is limited to 41 hours, which may affect model generalization. Baseline tests were conducted on only 10% of the dataset, which may not be comprehensive.
Plain Language Accessible to non-experts
Imagine a large library with various books representing different dialects and speakers. The Ramsa corpus is like this library, offering rich Emirati Arabic speech data. Researchers use this data to train and test speech recognition and synthesis models, much like librarians recommend books based on different genres. With this corpus, researchers can better understand and handle the diversity of Emirati Arabic.
ELI14 Explained like you're 14
Imagine you're in a school with classmates from different places, speaking different dialects. The Ramsa corpus is like a big classroom recording these conversations. Scientists use these recordings to teach computers how to understand and speak Emirati Arabic. Just like you need to listen and practice a new language, computers need this data to improve their language skills. In the future, computers might become smarter and understand more dialects and accents!
Glossary
ASR (Automatic Speech Recognition)
ASR refers to the process of converting spoken language into text by computer systems.
In this paper, ASR is used to evaluate the recognition performance of the corpus.
TTS (Text-to-Speech)
TTS is a technology that converts text into spoken voice, enabling computers to 'speak'.
TTS is used to test the corpus's performance in speech synthesis.
Whisper-large-v3-turbo
An open-source Transformer model for speech recognition with fast inference capabilities.
Performed best in ASR baseline testing.
MMS-TTS-Ara
A TTS model for Arabic, focusing on multilingual synthesis.
Performed best in TTS baseline testing.
Zero-shot setting
Testing a model's ability without specific training data.
Used to evaluate initial performance of ASR and TTS models.
Open Questions Unanswered questions from this research
- 1 How to improve ASR and TTS model performance on larger corpus scales?
- 2 How to increase representation of Mountain/Shihhi dialect in the corpus?
Applications
Immediate Applications
Speech Recognition Systems
Develop more accurate Emirati Arabic speech recognition systems for customer service and voice assistants.
Long-term Vision
Multilingual Education Platform
Create a platform supporting multiple dialects to help learners better master Arabic.
Abstract
Ramsa is a developing 41-hour speech corpus of Emirati Arabic designed to support sociolinguistic research and low-resource language technologies. It contains recordings from structured interviews with native speakers and episodes from national television shows. The corpus features 157 speakers (59 female, 98 male), spans subdialects such as Urban, Bedouin, and Mountain/Shihhi, and covers topics such as cultural heritage, agriculture and sustainability, daily life, professional trajectories, and architecture. It consists of 91 monologic and 79 dialogic recordings, varying in length and recording conditions. A 10\% subset was used to evaluate commercial and open-source models for automatic speech recognition (ASR) and text-to-speech (TTS) in a zero-shot setting to establish initial baselines. Whisper-large-v3-turbo achieved the best ASR performance, with average word and character error rates of 0.268 and 0.144, respectively. MMS-TTS-Ara reported the best mean word and character rates of 0.285 and 0.081, respectively, for TTS. These baselines are competitive but leave substantial room for improvement. The paper highlights the challenges encountered and provides directions for future work.