cs.LG 2607.24900

Inverse RL Helps Align AI by Imitating Humans

PARED employs inverse RL in feature space to extract implicit rewards from demonstrations, improving language model alignment without task-specific annotations.

Michał Wiliński, Liu Leqi, Chirag Nagpal

2026-07-28 42
cs.LG 2607.21542

Zero-Flow Two-Sample Tests

Proposes Zero-Flow Two-Sample Test (ZF2ST) with strong performance on synthetic and image datasets.

Yakun Wang, Leyang Wang, Song Liu et al.

2026-07-24 1
cs.LG 2607.21427

Context-weighted Discrete Flow Matching

Proposes context-weighted discrete flow matching, reducing perplexity by 63% via local context-aware sampling and loss.

Daniil Cherniavskii, Daniel Severo, Karen Ullrich

2026-07-23 23