The Fairness Collapse Phenomenon: Bias Amplification in Language Models Trained on Synthetic Data
Using Qwen2.5-0.5B, recursive training raised EO GAP 13.18→19.38 while PPL improved 16.07→10.55.
Irina Proskurina, Antoine Gourru, Julien Velcin
Using Qwen2.5-0.5B, recursive training raised EO GAP 13.18→19.38 while PPL improved 16.07→10.55.
Irina Proskurina, Antoine Gourru, Julien Velcin
DyLaR enhances video QA accuracy to 58.2% with under 20 tokens per query using dynamic latent reasoning.
Haotian Xia, Zilin Xiao, Junbo Zou et al.
Proposed FinProBench with RGRC extracts role-based standards from real deliverables, boosting financial AI evaluation efficiency.
Ben Wang, Kang Zhou, Lifan Guo et al.
ASP uses membrane potential as a Bayesian belief, enabling adaptive region selection for 3D point cloud recognition, achieving 90.62% accuracy with linear energy savings.
Akarsh Jain, Arya Pawa, Ayush Debnath et al.
Omega-S uses graph topology metrics to regularize LLM fine-tuning, improving memory retention by controlling weight connectivity diversity.
Alberto Acedo
Developed a multimodal robot interaction framework supporting inclusive wellbeing assessment for children with DLD and migration backgrounds, with ethical design principles.
Fethiye Irmak Dogan, Yue Lou, Alva Markelius et al.
AgenticVAU employs multi-agent explore-verify reasoning, surpassing zero-shot and RL baselines in video anomaly understanding with significant accuracy gains.
Yuxiang Duan, Huining Li, Ao Li et al.
Proposes black-box dispersion–revision diagnostic using CI to verify if output diversity correlates with genuine epistemic revision.
Molood Arman
VetScore integrates citation support and harm potential scoring for veterinary long-form QA verification, improving trustworthiness.
Ivan Kartáč, Jan Tovarys, Mateusz Lango et al.
CILER models latent environments using user-conditioned exponential families to enhance OOD recommendation performance.
Qianqian Wang, Wenwu Gong, Yunshan Li et al.
SEER exposes query-specific evidence for frozen VLMs, improving frozen GQA-Train900 accuracy by 3.94 points over Full.
Feixiang Liu, Likun Wang, Qiang Qiu et al.
SALT uses subspace-aligned training to recover high-rank accuracy with ultra-low-rank residuals, achieving 18.5% accuracy improvement.
Xiang Li, Pengcheng Wang, Huazheng Wang et al.
This study analyzes SFT and RL in multi-task learning, revealing RL induces sparse, orthogonal parameter updates, reducing task interference.
Kejian Zhu, Zhuoran Jin, Shangqing Tu et al.
Test-time augmentation improves robustness of tabular-to-image classifiers under distribution shifts; composite strategies perform best.
Malena Loza, Felipe Grijalva, Eva Milara et al.
Proposes KARAT heterogeneous system with general-purpose near-memory processing, boosting sparse attention throughput by up to 6.13× for long-context LLMs.
Hyungkyu Ham, Junhyeong Bae, Seungheon Lee et al.
LLaDA MoE v2 scales mixture-of-experts diffusion models, trained on 23.5T tokens, nearing Qwen3 performance.
Fengqi Zhu, Shaoxuan Xu, Jingyang Ou et al.
Proposes Approximate Speculative Decoding (ASD), a zero-training verifier that boosts throughput by up to 15.26% via budgeted longest-prefix selection.
Yuannuo Feng, Zegang Peng, Yuxin Xie et al.
Proposes CSP, a complex-valued state-space model relying solely on state propagation, achieving 100% accuracy on deterministic tasks.
Xiaohe Li, Yang Lu
Proposes SGFormer with Triple-Structure-Attention, significantly reducing attention divergence and boosting local feature matching accuracy by 2% on MegaDepth-1500.
Runyu Zhu
Proposes self-evolving coding agents leveraging executable feedback to continuously update behaviors and components, improving code quality and robustness.
Hao Zhou, Haichuan Hu, Ye Shang et al.