Hypernetworks for Perspectivist Adaptation
Utilizing hypernetwork and adapters architecture, this study enhances perspective adaptation in hate speech detection with fewer parameters.
Daniil Ignatev, Denis Paperno, Massimo Poesio
Utilizing hypernetwork and adapters architecture, this study enhances perspective adaptation in hate speech detection with fewer parameters.
Daniil Ignatev, Denis Paperno, Massimo Poesio
ExtremBench benchmark reveals LLM discrepancies in solving extremal problems.
Binxin Gao, Jingjun Han
GAR improves theorem proving efficiency via adversarial RL, achieving a 4.20% gain on MiniF2F-Test.
Ruida Wang, Jiarui Yao, Rui Pan et al.
Proposes end-to-end conformal risk training extending CRC to OCE risks, improving model performance and guarantees.
Christopher Yeh, Nicolas Christianson, Adam Wierman et al.
Proposes a dynamic optimal transport-based framework for multivariate counterfactual identification, ensuring uniqueness and consistency.
Fabio De Sousa Ribeiro, Ainkaran Santhirasekaram, Ben Glocker
GRADE framework uses group-relative policy optimization and Dirichlet exploration for personalized multi-task fusion, improving CTR by 0.595% and CVR by 1.193%.
Tingfeng Hong, Pingye Ren, Xinlong Xiao et al.
Proposes ETD, training models to iterate over key layers during mid-training, boosting reasoning accuracy by up to 36% on benchmarks.
Yeskendir Koishekenov, Aldo Lipani, Nicola Cancedda
Introduced Control-Augmented Autoregressive Diffusion (CADA) for data assimilation, significantly improving stability and accuracy.
Prakhar Srivastava, Farrin Marouf Sofian, Francesco Immorlano et al.
LMCACHE employs batch data movement and modular APIs to enable efficient cross-engine and hierarchical KV cache management for enterprise-scale LLM inference.
Yuhan Liu, Yihua Cheng, Jiayi Yao et al.
Derives closed-form Busemann functions in Wasserstein space, enabling efficient distribution slicing and distance computation.
Clément Bonet, Elsa Cazelles, Lucas Drumetz et al.
PolyKAN framework uses polyhedral analysis for KAN compression, offering theoretical guarantees on model reduction and error control.
Di Zhang
TokenFlow employs preemptive scheduling and proactive KV cache management to boost streaming throughput by 82.5%, reducing P99 TTFT by 80.2%.
Junyi Chen, Chuheng Du, Renyuan Liu et al.
EqM learns a time-invariant energy gradient landscape, enabling optimization-based sampling with FID 1.90, outperforming diffusion/flow models.
Runqian Wang, Yilun Du
Diffusion^2 uses 3D point clouds to generate RF heatmaps with just 1.9 dB error, 27x faster.
Kyoungjun Park, Yifan Yang, Changhan Ge et al.
This study demonstrates that a standard Transformer, trained directly on Cartesian coordinates without graph priors, can achieve energy and force prediction accuracy comparable to state-of-the-art equivariant GNNs on OMol25, with faster inference.
Tobias Kreiman, Yutong Bai, Fadi Atieh et al.
Unified protocol shows synthetic data can accelerate LLM pre-training; optimal mix around 30% synthetic improves efficiency 5-10x.
Feiyang Kang, Newsha Ardalani, Michael Kuchnik et al.
The study explores the generalization of preference optimization under noisy feedback, proposing the GPO method and validating its effectiveness.
Shawn Im, Sharon Li
This study introduces a systematic multi-turn RL framework for large language models, emphasizing environment, reward, and policy pillars, validated across TextWorld, ALFWorld, and SWE-Gym with key improvements of up to 88%.
Ruiyi Wang, Prithviraj Ammanabrolu
This paper provides the first theoretical analysis of one-layer Mamba's training dynamics and robustness to outliers in in-context learning.
Hongkang Li, Songtao Lu, Xiaodong Cui et al.
CRM links step rewards to outcomes via conditional hazards, reaching 43.3% on AIME24 in verifier-free RL, 16.7 points above PURE.
Zheng Zhang, Ziwei Shan, Kaitao Song et al.