Score Centering Stabilizes Off-policy Reinforcement Learning
Score Centering stabilizes off-policy RL by reducing training-inference mismatch drift.
Martin Marek, Max Ryabinin
Score Centering stabilizes off-policy RL by reducing training-inference mismatch drift.
Martin Marek, Max Ryabinin
PosteriorBench evaluates distributional accuracy of generative inverse solvers, revealing neural operators improve resolution robustness.
Jiachen Yao, Zi-Siang Hsu, Xi Deng et al.
Using 1D CNN for RF fingerprinting under heterogeneous protocol interference, achieving up to 97% accuracy.
Tariq Abdul-Quddoos, Xiangfang Li, Lijun Qian
ActObs method supervises observations to change exploration in RL, enhancing Qwen3 model performance on Terminal-Bench 2.0.
Juzheng Zhang, Disha Makhija, Manoj Ghuhan Arivazhagan et al.
Using the DAGMA algorithm, this paper identifies causal graphs and verifies edge direction identifiability in mixed datasets.
Sambit Mishra, Yingying Wang, Christine K. Johnson et al.
The study compares parallelism in three diffusion language models, finding Gaussian and uniform diffusion superior under certain conditions.
Sitan Chen, Liye Wang
Analyzes bias in nonlinear two-time-scale stochastic approximation under constant step-sizes, providing upper bounds on mean-squared error.
Djamel Rassem Lamouri, Dorian Baudry, Nicolas Gast
VAST method enhances accuracy in cold-start semi-supervised learning by geometric inference, outperforming existing baselines.
Itai David, Daphna Weinshall
Physical-State-Guided Diffusion Sampling (PSG) achieves superior geological structure recovery on OpenFWI datasets.
Chen Min, Haowen Jiang, Zheng Ma et al.
Using IQP circuits, Logistic Regression's F1 improves from 0.462 to 0.517, outperforming Kernel PCA.
Menachem Finkelstein, Diana Legziel Levy, Zohar Yakhini et al.
Introduces covariance neural networks (VNNs), combining PCA and graph learning to improve stability and transferability on multiscale datasets.
Saurabh Sihag, Andrea Cavallo, Elvin Isufi et al.
Introduced Semigroup-JEPA for zero-shot physics generalization, achieving 34% error reduction and 23.3% control success improvement.
Andy Zeyi Liu, Haoran Sun, Lucas Baker et al.
FOM-UL achieves efficient forgetting by selectively updating Transformer layers, maintaining utility and privacy under 8/4-bit quantization.
Ravi Ranjan, Olivera Kotevska, Agoritsa Polyzou
Kalman Delta Networks improve linear attention by modeling uncertainty, enhancing perplexity and accuracy in 750M and 1.3B parameter models.
Ngoc Bui, Tinglin Huang, Rex Ying
Online Draft Co-Training with Zigzag Ring Attention and TapChannel achieves up to 1.88× end-to-end RL speedup.
Zili Wang, Zhaopeng Qiu, Yuekai Zhang et al.
Introduced a new theoretical framework for Masked Pretraining (MPT) to address dimensional collapse.
Qi Zhang, Runyu Zhou, Yifei Wang et al.
RegionFed achieves personalized query understanding with gradient conflict analysis, reaching 92.27% accuracy.
Quoc H. Nguyen, Ali Lafzi, Abhijeet Phatak et al.
Introduces a two-level framework using reasoning distillation and product-type test-time training to enhance scalable recommendation, achieving AUC of 0.924.
Siliang Liu, Mohammad Ghasemi, Sapan Patel et al.
Introduced a Hessian-based variational continuation method to automate finding periodic orbits in double pendulum systems.
Leo Yao, Ziming Liu, Max Tegmark
EGF generates categorical graphs via continuous embeddings, achieving 0.150 FCD on QM9.
Ethan Ma, Zihan Wang, Chris Siu Yeung Chow et al.