Adversarial Training for Process Reward Models
APRM improves mathematical reasoning accuracy by 3.4 percentage points through adversarial training between a generator and reward model.
Gurusha Juneja, Deepak Nathani, William Yang Wang
APRM improves mathematical reasoning accuracy by 3.4 percentage points through adversarial training between a generator and reward model.
Gurusha Juneja, Deepak Nathani, William Yang Wang
Proposes TokCom-UEP, integrating non-uniform semantic importance with rateless UEP coding, boosting image transmission robustness.
Kaizheng Zhang, Zuolin Jin, Zhihang Cheng et al.
DeepSeekMath-V2 trains a faithful verifier and proof generator with self-verification, achieving high accuracy in mathematical reasoning.
Zhihong Shao, Yuxiang Luo, Chengda Lu et al.
CoT4AD enhances vision-language models in autonomous driving using Chain-of-Thought reasoning, achieving state-of-the-art performance.
Zhaohui Wang, Tengbo Yu, Hao Tang
Proposes SPLADE and Extended-SPLADE models with pruning strategies for billion-scale web document retrieval, balancing effectiveness and efficiency.
Taeryun Won, Tae Kwan Lee, Hiun Kim et al.
S2KAN integrates symbolic primitives with differentiable sparsity, achieving compact, interpretable models with high accuracy.
James Bagrow, Josh Bongard
Decentralized peer-to-peer framework Matrix boosts multi-agent synthetic data generation throughput by 2-15× via message-driven asynchronous scheduling.
Dong Wang, Yang Li, Ansong Ni et al.
EvilGenie benchmarks coding-agent reward hacking with holdouts, file monitoring, and LLM judges; GPT-5 had one false positive on unambiguous cases.
Jonathan Gabor, Jayson Lynch, Jonathan Rosenfeld
Task Arithmetic is the only method that reliably improves LLM performance in 'in-the-wild' settings.
Oğuz Kağan Hitit, Leander Girrbach, Zeynep Akata
Proposed STVG-o1 framework uses reinforcement fine-tuning with chain-of-thought bounding box reasoning, improving HCSTVG-v1 m tIoU by 7.3%.
Xin Gu, Haoji Zhang, Qihang Fan et al.
Proposed TALES taxonomy for evaluating cultural biases in LLM stories; 88% stories contain biases, especially in low-resource languages.
Kirti Bhagat, Shaily Bhatt, Athul Velagapudi et al.
Physics-informed Gaussian process prior for 2D stochastic Navier-Stokes based on quasi-Gaussianity theorem, capturing turbulence spectra.
Boumediene Hamzi, Houman Owhadi
Proposed retrieval-aware PMP enhances sarcasm detection, achieving 9.87% macro-F1 improvement on Twitter Indonesia dataset.
Michael Iskandardinata, William Christian, Derwin Suhartono
Evo-Memory evaluates self-evolving memory in LLMs during test-time, achieving an average performance gain of 0.65 across diverse tasks.
Tianxin Wei, Noveen Sachdeva, Benjamin Coleman et al.
Combines linear prediction, nearest neighbor classification, and conformal calibration for real-time flight safety risk warning with statistical guarantees.
Aaron O. Feldman, D. Isaiah Harp, Joseph Duncan et al.
∞-RoPE uses block-relative rotation and KV flush to enable infinite-length, controllable video generation.
Hidir Yesiltepe, Tuna Han Salih Meral, Adil Kaan Akan et al.
iMontage adapts pre-trained video models for multi-input/output image generation with enhanced dynamic range.
Zhoujie Fu, Xianfang Zeng, Jinghong Lan et al.
ReaDe employs a 'reason-then-describe' framework with two-stage training, achieving significant improvements in instruction fidelity and controllable video generation.
Shengqiong Wu, Weicai Ye, Yuanxing Zhang et al.
Proposes HHFT, a hierarchical heterogeneous feature Transformer, achieving +0.4% CTR AUC improvement and +0.6% GMV uplift on Taobao platform.
Liren Yu, Wenming Zhang, Silu Zhou et al.
DinoLizer leverages DINOv2 with LoRA fine-tuning to detect VAE and diffusion artifacts, achieving 20% higher IoU on forgery localization.
Minh Thong Doi, Vincent Itier, Jan Butora et al.