Pixels to Principles: Probing Intuitive Physics Understanding in Multimodal Language Models
Evaluated MLLMs using GRASP and IntPhys 2 datasets, found poor integration of visual and linguistic information.
Mohamad Ballout, Serwan Jassim, Elia Bruni
Evaluated MLLMs using GRASP and IntPhys 2 datasets, found poor integration of visual and linguistic information.
Mohamad Ballout, Serwan Jassim, Elia Bruni
Introduces a deterministic, modular framework with verifiable guard code for policy enforcement in LLM agent workflows, validated on τ-bench Airlines with 75% success.
Naama Zwerdling, David Boaz, Ella Rabinovich et al.
QuestA introduces partial solutions during RL training, boosting 1.5B models' math reasoning by over 10% on benchmarks.
Jiazheng Li, Hongzhou Lin, Hong Lu et al.
Infherno employs multi-step LLM agents with external tools to convert unstructured clinical notes into FHIR-compliant structured resources.
Johann Frei, Nils Feldhus, Lisa Raithel et al.
PoliAnalyzer uses NLP for personalized privacy policy analysis, achieving F1 scores of 90-100%.
Rui Zhao, Vladyslav Melnychuk, Jun Zhao et al.
MoR introduces dynamic recursive depths within shared Transformer layers, reducing validation perplexity by up to 15% and improving few-shot accuracy, while enhancing efficiency.
Sangmin Bae, Yujin Kim, Reza Bayat et al.
Introduces MisAttribution Framework and AttriData dataset; trains MisAttributionLLM for multi-task error evaluation with high correlation (0.935).
Zishan Xu, Shuyi Xie, Qingsong Lv et al.
MIRIX is a multi-agent, multimodal memory system that improves LLM accuracy by 35% and reduces storage by 99.9%.
Yu Wang, Xi Chen
SynthEHR-Eviction combines LLM-based synthetic data and prompt optimization to detect eviction statuses, achieving 88.8% Macro-F1.
Zonghai Yao, Youxia Zhao, Avijit Mitra et al.
Expanded MorphScore for 70 languages shows negligible correlation between morphological alignment and language model performance.
Catherine Arnett, Marisa Hudspeth, Brendan O'Connor
Gemini 2.5 Pro achieves SOTA in long multimodal reasoning with 3-hour video support using advanced Transformer architecture.
Gheorghe Comanici, Eric Bieber, Mike Schaekermann et al.
MemOS introduces a system-level memory management framework for LLMs, significantly improving long-term memory, personalization, and knowledge consistency.
Zhiyu Li, Chenyang Xi, Chunyu Li et al.
Selecting diverse, challenging reasoning traces from a strong teacher model enhances student model reasoning, outperforming existing datasets by 3-5% on benchmarks.
Yang Li, Youssef Emad, Karthik Padthe et al.
This study compares MLM and CLM pretraining, finding MLM superior in downstream tasks, but CLM more data-efficient and stable.
Hippolyte Gisserot-Boukhlef, Nicolas Boizard, Manuel Faysse et al.
This study compares manual, pattern-based, and LLM-generated CQ, revealing LLMs' efficiency but need for refinement.
Reham Alharbi, Valentina Tamma, Terry R. Payne et al.
AlignEvoSkill enhances skill evolution via knowledge tags and task alignment, achieving a 34.7% improvement.
Dingzirui Wang, Xuanliang Zhang, Keyan Xu et al.
HERB benchmark reveals retrieval bottlenecks in RAG systems over heterogeneous enterprise data, scoring 32.96 on average.
Prafulla Kumar Choubey, Xiangyu Peng, Shilpa Bhagavath et al.
Defines scoring bias, identifies three types—Rubric order, Score ID, reference answer bias—and proposes a quantification framework.
Qingquan Li, Shaoyu Dou, Kailai Shao et al.
RCO framework trains critic models via refinement signals, improving critique quality and response refinement by 15% on average.
Tianshu Yu, Chao Xiang, Mingchuan Yang et al.
FineWeb2 pipeline enables automatic filtering and deduplication supporting over 1000 languages, improving multilingual model performance.
Guilherme Penedo, Hynek Kydlíček, Vinko Sabolčec et al.