eess.AS 2410.00037

Moshi: a speech-text foundation model for real-time dialogue

Moshi employs hierarchical Transformer and residual vector quantization for real-time full-duplex speech dialogue with 160ms latency, advancing end-to-end voice systems.

Alexandre Défossez, Laurent Mazaré, Manu Orsini et al.

2024-09-18 43
cs.CV 2409.11340

OmniGen: Unified Image Generation

OmniGen is a unified diffusion model supporting multi-task image generation, including text-to-image, editing, and subject-driven tasks, with simplified architecture and knowledge transfer.

Shitao Xiao, Yueze Wang, Junjie Zhou et al.

2024-09-18 45
cs.CV 2409.09196

Are Sparse Neural Networks Better Hard Sample Learners?

This study shows sparse neural networks (SNNs) can match or outperform dense models on hard samples, especially with limited data and high difficulty, using specific sparsification strategies.

Qiao Xiao, Boqian Wu, Lu Yin et al.

2024-09-14 48