Domination-Avoiding Learning Agents Cannot Collude
The study proves that 'Domination-Avoiding' learning agents do not collude in markets, including mean-based and internal regret-minimizing algorithms.
Noam Nisan, Emmanuel Zerah
The study proves that 'Domination-Avoiding' learning agents do not collude in markets, including mean-based and internal regret-minimizing algorithms.
Noam Nisan, Emmanuel Zerah
TrOPD employs trust-region strategies and multiple KL estimators to stabilize on-policy distillation, outperforming SOTA with +6.18 performance points.
Xingrun Xing, Haoqing Wang, Boyan Gao et al.
Proposed SVHalluc benchmark evaluates speech-vision hallucination, revealing poor cross-modal alignment in state-of-the-art models with near-random accuracy.
Chenshuang Zhang, Kyeong Seon Kim, Chengxin Liu et al.
BraveGuard uses open-source threat mining and trajectory supervision to boost agent safety detection from 38.79% to 82.38%.
Yunhao Feng, Xiaohu Du, Xinhao Deng et al.
Proposes an interactive video world modeling framework integrating diffusion, VAE, and memory, achieving real-time multi-modal control with 95% scene accuracy.
Jiuming Liu, Chaojun Ni, Mengmeng Liu et al.
SkillRevise improves LLM agent success rate from 36.05% to 61.63% via execution-conditioned skill revision.
Yuxuan Liu, Zhaochen Su, Lingyun Xie et al.
LeAP learns permutation-based gates and removes 3,600+ dimensions from a 12,000+ dimensional industrial model with zero business-metric degradation.
Yihong Huang, Chen Chu, Fei Chen et al.
Proposed a decision-focused on-policy learning method for contextual linear optimization with partial feedback, showing lower cumulative regret than baselines.
Wyame Benslimane, Tinghan Ye, Pascal Van Hentenryck et al.
DAG-MoE improves MoE models by structural aggregation, enhancing language model performance.
Jiarui Feng, Hanqing Zeng, Karish Grover et al.
Decoupled Residual Denoising Diffusion (DRDD) introduces a two-stage process—stochastic noise diffusion for domain harmonization and deterministic residual diffusion for semantic mapping—achieving unified, data-efficient image-to-image translation with superior performance.
Ziyue Lin, Jiahe Hou, Hongyu Xia et al.
Single-layer phase encoder-decoder system achieves nonlinear function approximation via interference, with error below 10^-8, no nonlinear materials needed.
Yuntian Wang, Alexander Chen, Md Sadman Sakib Rahman et al.
$τ_0$-WM is a unified video-action world model excelling in long-horizon robotic tasks.
Pengfei Zhou, Shengcong Chen, Di Chen et al.
The paper argues LLMs should learn personalized rather than aggregated preferences, citing social choice theory and experimental data.
Cristina Garbacea
MBench evaluates long-term memory in video world models via entity, environment, and causal dimensions with 12 sub-metrics, using real long videos and VLM.
Shengjun Zhang, Zhang Zhang, Simin Huang et al.
Proposes Wan2.2 video diffusion model compression via few-step distillation and low-bit quantization, achieving superior efficiency.
Jinyang Du, Shenghao Jin, Ziqian Xu et al.
Introduces pause-and-think dataset and a 4B model fine-tuned with structured reasoning, achieving 58% accuracy with 59× fewer parameters than state-of-the-art models.
Shivam Singh, Saptarshi Majumder, Pratik Prabhanjan Brahma et al.
CARGO is a training-free routing framework using response agreement and Bayesian early stopping to decide when local models can answer reliably, supporting adjustable cloud offloading ratios.
Evan Chen, Shiqiang Wang, Kevin S Chan et al.
CAFOSat employs deep learning and GradCAM localization to refine annotations, covering 45,000 image patches across 20 US states for infrastructure-aware CAFO mapping.
Oishee Bintey Hoque, Nibir Chandra Mandal, Mandy L Wilson et al.
GNMR controls stability in low-precision language model training by comparing gradient norms with historical means.
Boao Kong, Weichen Jia, Engao Zhang et al.
Proposes semantic space matching using SSL features with Sinkhorn divergence, reducing ImageNet FID by 39×, improving one-step generative quality.
Hugues Van Assel, Edward De Brouwer, Saeed Saremi et al.