排序: 最新 热门 引用
cs.CL 2502.12788

Commonsense Reasoning in Arab Culture

提出ArabCulture数据集,评估阿拉伯文化常识推理,32B参数模型表现不佳。

Abdelrahman Sadallah, Junior Cedric Tonga, Khalid Almubarak 等

2025-02-18 3
cs.LG 2502.12292

Independence Tests for Language Models

提出一种独立性检验方法,能有效识别语言模型是否独立训练,p值精确。

Sally Zhu, Ahmed Ahmed, Rohith Kuditipudi 等

2025-02-18 5
cs.CL 2502.11187

TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking

TituLLMs为孟加拉语设计的首个大规模预训练模型,包含1B和3B参数,基于扩展的Llama-3.2 tokenizer,采用37亿标记数据,显著优于多语种版本。

Shahriar Kabir Nahin, Rabindra Nath Nandi, Sagor Sarker 等

2025-02-17 60