Online Learning under Delayed Feedback
Proposes BOLD and QPM-D algorithms for online learning with delayed feedback; theoretical bounds show multiplicative regret in adversarial and additive in stochastic settings.
Pooria Joulani, András György, Csaba Szepesvári