CommunityBench: Benchmarking Community-Level Alignment across Diverse Groups and Tasks
CommunityBench evaluates community-level model alignment, revealing LLMs' limitations in modeling community preferences.
Jiayu Lin, Zhongyu Wei
CommunityBench evaluates community-level model alignment, revealing LLMs' limitations in modeling community preferences.
Jiayu Lin, Zhongyu Wei
Proposes a Nash Social Welfare-based fair recommendation algorithm balancing match rate and fairness.
Yoji Tomita, Tomohiko Yokoyama
Proposes coherence optimization as a unified framework for self-improvement, proving its equivalence to description-length regularization, and introduces Gibbs sampling for scalable optimization.
Tianyi Qiu, Ahmed Hani Ismail, Zhonghao He et al.
Introduces RTCE benchmark using lossless compression algorithms to evaluate LLMs' bidirectional code reasoning and invertibility, revealing fundamental limitations.
Nickil Maveli, Antonio Vergari, Shay B. Cohen
This paper redefines tokenization as a core model design decision, proposing a context-aware co-design framework with standardized evaluation metrics.
Sawsan Alqahtani, Mir Tafseer Nayeem, Md Tahmid Rahman Laskar et al.
WorldMind aligns agentic world models via experiential learning, achieving 48.0% success rate on EB-ALFRED.
Baochang Ren, Yunzhi Yao, Rui Sun et al.
This study demonstrates that leaking evaluation prompts and gold nuggets can inflate RAG system scores, highlighting vulnerabilities in LLM-based evaluation methods.
Laura Dietz, Bryan Li, Eugene Yang et al.
This review frames AI4Math around AlphaGeometry, FunSearch, and AlphaEvolve, arguing for AI that discovers insight beyond verification.
Haocheng Ju, Bin Dong
FastAV reduces AV-LLM inference computation by over 40% using a two-stage pruning strategy.
Chaeyoung Jung, Youngjoon Jang, Seungwoo Lee et al.
HyFormer unifies sequence modeling and feature interaction via global tokens, boosting CTR prediction by 3-5% over baselines.
Yunwen Huang, Shiyong Hong, Xijun Xiao et al.
The study compares standard and masked autoencoders for audio anomaly detection, finding masked autoencoders provide more precise explanations.
Maab Elrashid, Anthony Deschênes, Cem Subakan et al.
Introducing 'Information Farming', a paradigm shift from passive harvesting to active cultivation using generative AI, with a farming analogy and empirical validation.
Leif Azzopardi, Adam Roegiest
LongPAS method enhances long-context reasoning, significantly outperforming RLVR baselines.
Miao Peng, Weizhou Shen, Nuo Chen et al.
CTC-DID employs CTC loss for Arabic dialect identification, excelling in low-resource settings.
Muhammad Umar Farooq, Oscar Saz
EMoE introduces eigenbasis-guided routing, balancing expert utilization and diversity, achieving 88.14% Top-1 accuracy on ImageNet.
Anzhe Cheng, Shukai Duan, Shixuan Li et al.
Proposed a multi-agent pipeline to transform reviews into actionable business advice, enhancing actionability and relevance.
Kartikey Singh Bhandari, Tanish Jain, Archit Agrawal et al.
AVIR combines lightweight retrieval and adaptive filtering to reduce 70% pages, achieving 84.58% ANLS in multi-page VQA.
Zongmin Li, Yachuan Li, Lei Kang et al.
SeLop method uses low-rank orthogonal projection to remove spurious bias, improving face forgery detection generalization.
Chi Wang, Xinjue Hu, Boyu Wang et al.
Unified coverage framework Uρ interpolates between KL divergence, average, and minimax exploration via parameter ρ, optimizing state-action visitation in reward-free MDPs.
Xihe Gu, Urbashi Mitra, Tara Javidi
Proposes new long-context robust activation probes, improving misuse detection accuracy and efficiency.
János Kramár, Joshua Engels, Zheng Wang et al.