Advancing Arabic Speech Recognition Through Large-Scale Weakly Supervised Learning
Trained an Arabic ASR system using Conformer architecture and weak supervision, achieving superior results.
Mahmoud Salhab, Marwan Elghitany, Shameed Sait et al.
Trained an Arabic ASR system using Conformer architecture and weak supervision, achieving superior results.
Mahmoud Salhab, Marwan Elghitany, Shameed Sait et al.
WebRollback enhances web agents with explicit rollback mechanisms, achieving top performance on Mind2Web-Live and WebVoyager benchmarks.
Zhisong Zhang, Tianqing Fang, Kaixin Ma et al.
Developed a closure principle and the closed eBH method, significantly improving FDR control under arbitrary dependence structures.
Ziyu Xu, Lasse Fischer, Aaditya Ramdas
AskQE employs a QA framework with LLaMA-3 70B, achieving high correlation (τ=0.878) with human judgments for critical MT error detection.
Dayeon Ki, Kevin Duh, Marine Carpuat
DeepMath-103K is a large, high-difficulty, decontaminated, verifiable math dataset that advances reasoning models.
Zhiwei He, Tian Liang, Jiahao Xu et al.
PartField uses contrastive learning for 3D part features, achieving 85% accuracy and 20x speedup on ShapeNetPart.
Minghua Liu, Mikaela Angelina Uy, Donglai Xiang et al.
TRELAWNEY rearranges training sequences, raising Star Graph accuracy from 0.05 to 1.00 on G(20,5).
Abitha Thankaraj, Yiding Jiang, J. Zico Kolter et al.
Proposes Wasserstein barycenter-based collaborative Bayesian optimization, effectively preserving data privacy while maintaining high performance.
Donglin Zhan, Haoting Zhang, Rhonda Righter et al.
Proposes Energy Matching, unifying flow matching and energy-based models, achieving state-of-the-art fidelity on CIFAR-10 and ImageNet.
Michal Balcerak, Tamaz Amiranashvili, Antonio Terpin et al.
GUI-R1 employs rule-based reinforcement fine-tuning with only 3K samples, outperforming SOTA on multi-platform GUI tasks.
Run Luo, Lu Wang, Wanwei He et al.
Introduces LLM-SRBench with 239 problems, evaluating LLMs' scientific equation discovery, emphasizing reasoning beyond memorization.
Parshin Shojaee, Ngoc-Hieu Nguyen, Kazem Meidani et al.
Proposes a cultural localization benchmark; finds explicit prompts boost cultural responses but reduce diversity; uses activation steering to control cross-lingual cultural responses.
Veniamin Veselovsky, Berke Argin, Benedikt Stroebl et al.
SocioVerse employs large-scale real user data and LLM agents to simulate diverse social scenarios with 84.3% accuracy in election prediction.
Xinnong Zhang, Jiayu Lin, Xinyi Mou et al.
SegEarth-R1 leverages large language models for geospatial pixel reasoning, achieving 70.75% cIoU on EarthReason, outperforming prior methods.
Kaiyu Li, Zepeng Xin, Li Pang et al.
Introduces SIFT-50M, a 50-million-sample multilingual speech instruction dataset, enabling models to outperform existing benchmarks.
Prabhat Pandey, Rupak Vignesh Swaminathan, K V Vijay Girish et al.
GigaTok scales visual tokenizers to 3 billion parameters using semantic regularization, resolving the reconstruction-generation trade-off and achieving state-of-the-art results.
Tianwei Xiong, Jun Hao Liew, Zilong Huang et al.
KARMMA framework achieves robust egocentric action recognition under missing modalities, reducing computational resources by 50%.
Maria Santos-Villafranca, Dustin Carrión-Ojeda, Alejandro Perez-Yus et al.
Unified framework for diverse spoken language models (SLMs), including pure speech, speech+text, and speech-aware text models, advancing universal speech processing.
Siddhant Arora, Kai-Wei Chang, Chung-Ming Chien et al.
Proposes a hybrid AI router-based super agent system combining local and cloud models for efficient task routing.
Yuhang Yao, Haixin Wang, Yibo Chen et al.
The AI Scientist-v2 employs agentic tree search, multi-modal feedback, and autonomous experiment management to generate peer-review-accepted scientific papers, advancing AI-driven research automation.
Yutaro Yamada, Robert Tjarko Lange, Cong Lu et al.