SODA: Semi On-Policy Black-Box Distillation for Large Language Models
SODA introduces semi on-policy black-box distillation, achieving 10× faster training with comparable or better performance.
Xiwen Chen, Jingjing Wang, Wenhui Zhu et al.
SODA introduces semi on-policy black-box distillation, achieving 10× faster training with comparable or better performance.
Xiwen Chen, Jingjing Wang, Wenhui Zhu et al.
An end-to-end framework combining natural language reasoning and formal verification to automate research-level mathematical proofs, demonstrated on open problems with minimal human input.
Haocheng Ju, Guoxiong Gao, Jiedong Jiang et al.
CountsDiff models natural number distributions with a novel p(t) schedule and non-monotonic reverse dynamics, excelling in image and scRNA-seq data imputation.
Renzo G. Soatto, Anders Hoel, Greycen Ren et al.
Hierarchical Planning with Latent World Models (HWM) enables zero-shot long-horizon visual control, achieving 70% success on real robot tasks, with 3× less computation.
Wancong Zhang, Basile Terver, Artem Zholus et al.
CGCMA uses conditionally-gated cross-modal attention to achieve event-conditioned asynchronous fusion, improving Sharpe ratio to +0.449 on CryptoMI dataset.
Yunxiang Guo
HabitatAgent achieves 95% accuracy in housing consultation using a multi-agent system.
Hongyang Yang, Yanxin Zhang, Yang She et al.
HISA employs a hierarchical index to accelerate sparse attention, reducing complexity from O(L^2) to near O(L) without retraining, achieving up to 3.75× speedup at 64K context.
Yufei Xu, Fanxu Meng, Fan Jiang et al.
DBR-AF framework achieves multivariate time series anomaly detection via dual-branch reconstruction and autoregressive flow, outperforming existing methods.
Jun Liu, Ying Chen, Ziqian Lu et al.
GIFT uses geometric feedback for self-bootstrapping, boosting IoU by 12% and reducing inference costs by 80% in image-to-CAD synthesis.
Giorgio Giannone, Anna Clare Doris, Amin Heyrani Nobari et al.
Proposes a curvature-aware Expected Free Energy acquisition function for Bayesian optimization, outperforming state-of-the-art methods in joint learning and optimization tasks.
Ajith Anil Meera, Wouter Kouw
VAN-AD combines visual MAE with normalizing flow for enhanced time series anomaly detection.
PengYu Chen, Shang Wan, Xiaohou Shi et al.
DataFlex unifies data selection, mixing, and reweighting, boosting large model training efficiency and accuracy.
Hao Liang, Zhengyang Zhao, Meiyi Qiang et al.
Proposed ARTA enhances multivariate time-series anomaly detection robustness via sparsity-constrained adversarial perturbations, outperforming SOTA benchmarks.
Hadi Hojjati, Narges Armanfard
Multi-Answer RL trains models to generate multiple plausible answers in one pass, improving diversity and calibration, with 50% accuracy boost on coding tasks.
Isha Puri, Mehul Damani, Idan Shenfeld et al.
PyHealth framework evaluates attention mechanisms in clinical time-series models; Chefer method proves most effective and efficient.
Yongda Fan, John Wu, Andrea Fitzpatrick et al.
UI-Voyager employs RFT and GRSD, achieving 81% success on AndroidWorld with a 4B model, surpassing human performance.
Zichuan Lin, Feiyu Liu, Yijun Yang et al.
SortedRL accelerates RL training for LLMs through online length-aware scheduling, enhancing efficiency and performance.
Yiqi Zhang, Huiqiang Jiang, Xufang Luo et al.
Introduces Graph Energy Matching (GEM), surpassing discrete diffusion models in molecular graph generation.
Michal Balcerak, Suprosana Shit, Chinmay Prabhakar et al.
Proposes DAK-UCB, integrating diversity and fidelity for online generative model selection using kernel UCB.
Donya Jafari, Farzan Farnia
Scaling DoRA achieves high-rank adaptation via factored norms and fused kernels, significantly reducing memory usage and enhancing speed.
Alexandra Zelenin, Alexandra Zhuravlyova