Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs
Sailor2 achieves a 50-50 win rate against GPT-4o in SEA languages, supporting 13 languages.
Longxu Dou, Qian Liu, Fan Zhou et al.
Sailor2 achieves a 50-50 win rate against GPT-4o in SEA languages, supporting 13 languages.
Longxu Dou, Qian Liu, Fan Zhou et al.
RG-VFM extends variational flow matching to Riemannian manifolds, improving protein/material generation by capturing curvature effects.
Olga Zaghen, Floor Eijkelboom, Alison Pouplin et al.
Proposes MeCo, a meta-cognition-based adaptive tool-use method, improving decision accuracy by 10-15% and reducing unnecessary calls by 20-25%.
Wenjun Li, Dexun Li, Kuicai Dong et al.
Introduced ArabCulture dataset to evaluate Arabic cultural commonsense reasoning; 32B models struggle.
Abdelrahman Sadallah, Junior Cedric Tonga, Khalid Almubarak et al.
IREGB achieves asymptotically optimal MIR exploration in O(K log K) under stochastically ordered priors.
Gal Bahar, Omer Ben-Porat, Kevin Leyton-Brown et al.
Expanding the V-Model for AI complex systems, achieving 30% safety improvement in autonomous driving.
Lars Ullrich, Michael Buchholz, Klaus Dietmayer et al.
UXAgent employs large language model (LLM) agents to simulate diverse user behaviors for web usability testing, providing rich qualitative and quantitative data.
Yuxuan Lu, Bingsheng Yao, Hansu Gu et al.
YOLOv12 integrates attention mechanisms with efficient architecture, achieving 40.6% mAP at 1.64ms on T4 GPU, outperforming YOLOv10/11.
Yunjie Tian, Qixiang Ye, David Doermann
The study reviews video-to-music generation techniques, focusing on conditioning input construction, conditioning mechanism, and music generation frameworks.
Shulei Ji, Songruoyao Wu, Zihao Wang et al.
Proposed an independence test method that effectively identifies if language models are independently trained, with precise p-values.
Sally Zhu, Ahmed Ahmed, Rohith Kuditipudi et al.
SWE-Lancer benchmark evaluates 1,400 real freelance software tasks; current models achieve under 45% success, far from earning $1 million.
Samuel Miserendino, Michele Wang, Tejal Patwardhan et al.
Introduced UTR framework to address temporal hacking in video MLLMs, enhancing video comprehension.
En Yu, Kangheng Lin, Liang Zhao et al.
This study evaluates LLMs on real-world mathematical definitions, introduces Def_Wiki and Def_ArXiv datasets, and enhances performance via feedback and grounding strategies.
Lan Zhang, Marco Valentino, Andre Freitas
Explores scaling laws for upscaling neural networks, revealing relationships between model size and performance.
Ayan Sengupta, Yash Goel, Tanmoy Chakraborty
Transferability-aware hypernet framework with H-embedding significantly improves continual learning by capturing task relationships, achieving top performance on benchmarks.
Yanru Wu, Jianning Wang, Xiangyu Chen et al.
Proposes DualIPW, combining query-level and position-level click propensity estimation, to mitigate relevance saturation bias in unbiased learning to rank.
Lulu Yu, Keping Bi, Jiafeng Guo et al.
OctoTools is a training-free, multi-agent framework with standardized tools, boosting accuracy by 9.3% across 16 complex reasoning tasks.
Pan Lu, Bowen Chen, Sheng Liu et al.
LLMs in medicine show promise but face challenges like hallucination management and multimodal integration.
Wenxuan Wang, Zizhan Ma, Zheng Wang et al.
TituLLMs, a pioneering Bangla LLM with 1B and 3B parameters, uses extended tokenizer and 37B token dataset, outperforming multilingual versions in key tasks.
Shahriar Kabir Nahin, Rabindra Nath Nandi, Sagor Sarker et al.
CMCTS framework enhances LLM's mathematical reasoning with constrained action space; 7B model achieves 83.4% accuracy.
Qingwen Lin, Boyan Xu, Guimin Hu et al.