RouterBench: A Benchmark for Multi-LLM Routing System
Introduces RouterBench, a benchmark with 405k samples, evaluating multi-LLM routing systems for cost and performance.
Qitian Jason Hu, Jacob Bieker, Xiuyu Li et al.
Introduces RouterBench, a benchmark with 405k samples, evaluating multi-LLM routing systems for cost and performance.
Qitian Jason Hu, Jacob Bieker, Xiuyu Li et al.
Proposed pseudo-point SPA algorithm significantly improves vertex hunting accuracy and speed.
Jiashun Jin, Zheng Tracy Ke, Gabriel Moryoussef et al.
Proposes Wukong, a stacking FM-based architecture, establishing a recommendation scaling law with performance beyond 100 GFLOP/sample.
Buyun Zhang, Liang Luo, Yuxin Chen et al.
Curiosity-driven red teaming (CRT) leverages exploration rewards to enhance test coverage, successfully eliciting toxic responses from LLaMA2, with a 19.6% toxicity rate compared to 10.2% by baseline methods.
Zhang-Wei Hong, Idan Shenfeld, Tsun-Hsuan Wang et al.
ArCHer employs hierarchical RL with high-level value functions guiding token-level policies, achieving 100x sample efficiency in multi-turn tasks.
Yifei Zhou, Andrea Zanette, Jiayi Pan et al.
Proposes HSTU architecture with 1.5 trillion parameters for generative recommendation, achieving 12.4% online metric improvement, handling sequences up to 8192 length, outperforming traditional Transformers.
Jiaqi Zhai, Lucy Liao, Xing Liu et al.
Proposes QontOT, a quantum-based framework for contextual optimal transport, outperforming classical neural OT in predicting distribution variations.
Nicola Mariella, Albert Akhriev, Francesco Tacchino et al.
MobileLLM employs deep-thin transformer architecture with parameter sharing, achieving state-of-the-art accuracy for sub-billion models on mobile devices.
Zechun Liu, Changsheng Zhao, Forrest Iandola et al.
Proposes VPDD, a discrete diffusion-based framework leveraging large-scale actionless human videos for multi-task robot policy transfer, outperforming state-of-the-art methods.
Haoran He, Chenjia Bai, Ling Pan et al.
Introducing Gibbs-based kernel herding with theoretical guarantees surpassing classical Monte Carlo.
Martin Rouault, Rémi Bardenet, Mylène Maïda
Proposes an exploration-based fair classification framework with theoretical guarantees on exploration, FDR control, and convergence to optimal classifiers.
Vijay Keswani, Anay Mehrotra, L. Elisa Celis
Introduced Rectify-Router method to enhance MoE model performance, achieving a 4.7% accuracy improvement.
Zhiyuan Zeng, Qipeng Guo, Zhaoye Fei et al.
TimeSeriesBench offers 168 evaluation settings to enhance the industrial applicability of time series anomaly detection models.
Haotian Si, Jianhui Li, Changhua Pei et al.
SpLiCE transforms CLIP embeddings into sparse, interpretable semantic concepts, maintaining high zero-shot accuracy.
Usha Bhalla, Alex Oesterling, Suraj Srinivas et al.
StrongREJECT benchmark uses high-quality forbidden prompts and an automated evaluator to accurately measure jailbreak effectiveness, reducing overestimation issues.
Alexandra Souly, Qingyuan Lu, Dillon Bowen et al.
Proposes a causal reasoning framework based on reinforcement learning, utilizing two key lemmas to quantify causality in multivariate stochastic processes.
Mehdi Fatemi, Sindhu Gowda
InfoRM introduces an information bottleneck for reward modeling, effectively mitigating reward hacking and improving generalization in RLHF.
Yuchun Miao, Sen Zhang, Liang Ding et al.
Proposes PNSIS framework, integrating necessity and sufficiency probabilities to extract invariant subgraphs, boosting graph OOD generalization.
Xuexin Chen, Ruichu Cai, Kaitao Zheng et al.
Introduces Target Score Matching (TSM), leveraging known target scores to improve low-noise score estimates, enhancing diffusion models’ accuracy.
Valentin De Bortoli, Michael Hutchinson, Peter Wirnsberger et al.
Introduces granular hyperparameter G and scaling laws for MoE, optimizing training to outperform dense Transformers under various compute budgets.
Jakub Krajewski, Jan Ludziejewski, Kamil Adamczewski et al.