NVV-Locator: From Transcript Tags to Acoustic Boundaries for Fine-Grained Nonverbal Vocalization Grounding
NVV-Locator uses a non-autoregressive slot-filling architecture to achieve 71.0% Micro F1 in precise NVV event localization.
Key Findings
Methodology
NVV-Locator employs a non-autoregressive slot-filling architecture, integrating an audio encoder and a language model decoder to predict lexical timestamps and NVV boundaries. It unifies 26 NVV categories, constructs a large-scale timestamp-supervised dataset, and introduces the NVV-TimeBench benchmark for evaluation.
Key Results
- On NVV-TimeBench, NVV-Locator achieves 71.0% Micro F1 and 70.2% Macro F1, with a Macro mIoU of 80.4% and a Macro mMAE of 59.6 ms, significantly outperforming existing large audio models.
- Cross-corpus generalization of NVV-Locator is validated on external corpora, demonstrating robustness across different datasets.
- Ablation studies confirm the advantages of the non-autoregressive slot-filling architecture in temporal consistency and boundary precision.
Significance
This research holds significant value for academia and industry, addressing shortcomings in existing methods for NVV recognition and temporal localization, and offering new possibilities for speech understanding and generation. Accurate temporal localization enables applications in audio editing and emotion analysis.
Technical Contribution
NVV-Locator advances beyond existing methods by achieving higher temporal localization accuracy through a non-autoregressive slot-filling architecture. It introduces a unified 26-category NVV taxonomy and a large-scale timestamp-supervised dataset, laying a new foundation for NVV research.
Novelty
NVV-Locator is the first to apply a non-autoregressive slot-filling architecture to NVV temporal localization, significantly enhancing temporal consistency and boundary precision compared to existing methods.
Limitations
- Processing long audio durations may lead to excessive computational resource consumption.
- Some NVV categories remain sparse in the dataset, potentially affecting recognition accuracy.
Future Work
Future research directions include expanding the dataset to cover more NVV categories and optimizing the model for efficient processing of long audio durations.
AI Executive Summary
Human speech includes nonverbal vocalizations (NVVs) such as laughter, sighs, and breaths, which convey affective and interactional information. However, existing methods typically represent these events as transcript-level tags, lacking precise waveform-time boundaries. NVV-Locator employs a non-autoregressive slot-filling architecture to achieve fine-grained NVV temporal grounding. This method integrates an audio encoder and a language model decoder to predict lexical timestamps and NVV boundaries. By unifying 26 NVV categories, constructing a large-scale timestamp-supervised dataset, and introducing the NVV-TimeBench benchmark, it sets a new standard for evaluation.
On NVV-TimeBench, NVV-Locator achieves 71.0% Micro F1 and 70.2% Macro F1, with a Macro mIoU of 80.4% and a Macro mMAE of 59.6 ms, significantly outperforming existing large audio models. Ablation studies confirm the advantages of the non-autoregressive slot-filling architecture in temporal consistency and boundary precision. This research holds significant value for academia and industry, addressing shortcomings in existing methods for NVV recognition and temporal localization, and offering new possibilities for speech understanding and generation.
Future research directions include expanding the dataset to cover more NVV categories and optimizing the model for efficient processing of long audio durations. The success of NVV-Locator demonstrates the potential of non-autoregressive slot-filling architectures in temporal localization tasks, advancing NVV research.
Deep Analysis
Background
Nonverbal vocalizations (NVVs) like laughter, sighs, and breaths convey emotional and interactional information in human communication. However, these events often lack stable linguistic forms, making them difficult to accurately recognize and localize in continuous speech. Existing methods typically represent NVVs as transcript-level tags, lacking precise waveform-time boundaries.
Core Problem
Existing methods fall short in NVV recognition and temporal localization, failing to provide precise waveform-time boundaries. This limits the capability of speech understanding and generation, especially in applications requiring fine-grained temporal localization.
Innovation
NVV-Locator achieves fine-grained NVV temporal grounding through a non-autoregressive slot-filling architecture. This method integrates an audio encoder and a language model decoder to predict lexical timestamps and NVV boundaries. By unifying 26 NVV categories, constructing a large-scale timestamp-supervised dataset, and introducing the NVV-TimeBench benchmark, it sets a new standard for evaluation.
Methodology
- �� Employs a non-autoregressive slot-filling architecture to avoid temporal inconsistency.
- �� Integrates an audio encoder and a language model decoder to predict lexical timestamps and NVV boundaries.
- �� Unifies 26 NVV categories and constructs a large-scale timestamp-supervised dataset.
- �� Introduces the NVV-TimeBench benchmark for evaluation.
Experiments
Experiments are conducted on NVV-TimeBench to evaluate NVV-Locator's performance. It is compared with four large audio models, including Gemini-2.5-Pro, Qwen3-Omni-Instruct, Step-Audio-R1.1, and MOSS-Audio-8B-Instruct. Evaluation metrics include Micro F1, Macro F1, Macro mIoU, and Macro mMAE.
Results
NVV-Locator achieves 71.0% Micro F1 and 70.2% Macro F1 on NVV-TimeBench, with a Macro mIoU of 80.4% and a Macro mMAE of 59.6 ms, significantly outperforming existing large audio models.
Applications
NVV-Locator can be used in applications such as audio editing and emotion analysis, especially in scenarios requiring fine-grained temporal localization. Its precise temporal localization capability offers new possibilities for speech understanding and generation.
Limitations & Outlook
Processing long audio durations may lead to excessive computational resource consumption. Some NVV categories remain sparse in the dataset, potentially affecting recognition accuracy. Future research directions include expanding the dataset and optimizing the model.
Plain Language Accessible to non-experts
Imagine you're at a concert, surrounded by people cheering, clapping, and even whistling. These sounds aren't part of the music but convey the audience's emotions and interactions. NVV-Locator is like a super microphone that can precisely capture these non-musical sounds and tell you their exact position in the music. This helps music producers better understand audience reactions and make more refined edits in post-production.
ELI14 Explained like you're 14
Imagine you're playing a music game, where there's not only music but also cheers and claps from the audience. These sounds make the game more fun, but you need to know their exact position to score points. NVV-Locator is like a super helper in the game, accurately telling you the start and end times of these sounds. This way, you can better master the game rhythm and score high! Isn't that cool?
Glossary
Nonverbal Vocalizations
Refers to sounds like laughter, sighs, and breaths that convey emotions without linguistic information.
Used to identify and localize these events in speech.
Slot-Filling Architecture
A model architecture for predicting timestamps and event categories, avoiding temporal inconsistency.
Core architecture used in NVV-Locator.
Timestamp-Supervised Dataset
A dataset with precise time annotations used to train models for temporal localization.
Key dataset for training NVV-Locator.
Macro F1
A metric evaluating a model's average performance across different categories.
Used to assess NVV-Locator's performance on NVV-TimeBench.
Macro mIoU
A metric evaluating the overlap between predicted and true boundaries.
Used to evaluate NVV-Locator's temporal localization accuracy.
Open Questions Unanswered questions from this research
- 1 How to effectively recognize and localize NVVs in long audio durations remains an open question.
- 2 Some NVV categories are sparse in the dataset, affecting recognition accuracy.
Applications
Immediate Applications
Audio Editing
By precisely localizing NVVs, it helps audio engineers make more refined edits in post-production.
Long-term Vision
Emotion Analysis
By recognizing and analyzing NVVs, it helps researchers better understand human emotions and interactions.
Abstract
Human speech includes nonverbal vocalizations (NVVs), such as laughter, sighs, breaths, and coughs, which convey affective and interactional information. Existing approaches typically represent NVVs as transcript-level tags, providing limited supervision for their waveform-time boundaries. We present NVV-Locator for fine-grained NVV temporal grounding. We first unify 26 NVV categories across public resources and construct large-scale timestamp-supervised training data through dual-LLM verification, transcript-guided forced alignment, and energy-based boundary refinement. We further introduce NVV-TimeBench, an expert-refined benchmark with 667 utterances and 1,094 events. NVV-Locator uses a non-autoregressive slot-filling architecture to jointly predict lexical timestamps, NVV categories, and event boundaries. On NVV-TimeBench, it achieves 71.0% Micro F1, 70.2% Macro F1, 80.4% Macro mIoU, and 59.6 ms Macro mMAE, outperforming the evaluated large audio model counterparts. Evaluation on an external corpus further demonstrates the cross-corpus generalization of NVV-Locator.