Choosing Online Experiment Designs under Interference in Ads, Recommendations, and Member-Experience Systems

TL;DR

The study proposes an interference-aware experiment design framework, selecting designs in ads and recommendation systems, with a risk of 1.295 on Criteo ads.

stat.ML 🔴 Advanced 2026-05-25 5 views
Prashant Shekhar Caroline Howard
experiment design ad systems recommendation systems interference mechanisms robustness

Key Findings

Methodology

The paper proposes an experiment design framework based on mechanism-robust design decisions, considering six implementable designs, compared by worst-case planning risk. Risk factors include exposure bias, assignment-unit variance, minimum detectable effect, contamination or carryover, operational cost, and estimand mismatch. The design bias is quantified using Wasserstein distance, providing geometric guarantees under Lipschitz exposure response.

Key Results

  • On the Criteo ads dataset, user randomization was selected with a dimensionless robust risk of 1.295.
  • On the Open Bandit bts/men dataset, switchback design was chosen with a risk of 2.105.
  • On the KuaiRand dataset, cluster randomization was selected with a risk of 2.240.

Significance

The study offers a new perspective on online experiment design in ads, recommendation, and member-experience systems, especially under uncertain interference mechanisms. By considering multiple possible interference mechanisms, the study proposes a method capable of making robust design choices under uncertainty. This approach is significant for improving the accuracy and reliability of experiment designs, especially in modern complex systems.

Technical Contribution

The technical contributions include a new experiment design selection framework capable of robust selection under uncertain interference mechanisms. By introducing Wasserstein distance to quantify design bias and providing geometric guarantees under Lipschitz exposure response, the proposed selector algorithm achieves exact recovery and excess-risk control within a finite design catalog.

Novelty

This study is the first to incorporate uncertainty in interference mechanisms into the experiment design selection problem, using a geometric approach based on Wasserstein distance to quantify design bias. Compared to existing work, this method provides robust design choices across multiple interference mechanisms.

Limitations

  • The method may be limited by the need for extensive historical data and product knowledge to construct the uncertainty set.
  • Design choices may depend on specific risk weight settings in some cases.

Future Work

Future work could explore applying this framework to larger design catalogs and better estimating uncertainty sets in practical applications.

AI Executive Summary

In ads, recommendation, and member-experience systems, online experiment designs are often planned before the dominant interference mechanism is known. Existing methods struggle to select appropriate designs under uncertain interference mechanisms. This paper proposes an experiment design framework based on mechanism-robust design decisions, considering six implementable designs, compared by worst-case planning risk. Risk factors include exposure bias, assignment-unit variance, minimum detectable effect, contamination or carryover, operational cost, and estimand mismatch. The design bias is quantified using Wasserstein distance, providing geometric guarantees under Lipschitz exposure response. Experimental results show that the selector gives different design recommendations across datasets. On the Criteo ads dataset, user randomization was selected with a dimensionless robust risk of 1.295; on the Open Bandit bts/men dataset, switchback design was chosen with a risk of 2.105; on the KuaiRand dataset, cluster randomization was selected with a risk of 2.240. The study offers a new perspective on online experiment design in ads, recommendation, and member-experience systems, especially under uncertain interference mechanisms. By considering multiple possible interference mechanisms, the study proposes a method capable of making robust design choices under uncertainty. This approach is significant for improving the accuracy and reliability of experiment designs, especially in modern complex systems. Future work could explore applying this framework to larger design catalogs and better estimating uncertainty sets in practical applications.

Deep Analysis

Background

In ads, recommendation, and member-experience systems, online experiment design is central to product and machine-learning decisions. However, the complexity of interference mechanisms in modern systems makes experiment design challenging. Existing methods often assume known interference mechanisms, but in practice, these mechanisms are often uncertain.

Core Problem

The core problem is selecting appropriate experiment designs under uncertain interference mechanisms. Interference mechanisms may propagate through budgets, inventory, producer exposure, graph spillovers, or temporal carryover, making randomization design itself a statistical decision.

Innovation

The core innovation is an experiment design framework based on mechanism-robust design decisions, considering multiple possible interference mechanisms to make robust design choices under uncertainty.

Methodology

  • �� Proposes a new experiment design selection framework capable of robust selection under uncertain interference mechanisms.
  • �� Introduces Wasserstein distance to quantify design bias, providing geometric guarantees under Lipschitz exposure response.
  • �� The selector algorithm achieves exact recovery and excess-risk control within a finite design catalog.

Experiments

Experiments were conducted using public datasets like Criteo, Open Bandit, and KuaiRand. The selector was tested on these datasets to evaluate different designs' performance under various interference mechanisms.

Results

Experimental results show that the selector gives different design recommendations across datasets. On the Criteo ads dataset, user randomization was selected with a dimensionless robust risk of 1.295; on the Open Bandit bts/men dataset, switchback design was chosen with a risk of 2.105; on the KuaiRand dataset, cluster randomization was selected with a risk of 2.240.

Applications

The framework can be directly applied to online experiment design in ads, recommendation, and member-experience systems, especially under uncertain interference mechanisms.

Limitations & Outlook

The method may be limited by the need for extensive historical data and product knowledge to construct the uncertainty set. Design choices may depend on specific risk weight settings in some cases.

Plain Language Accessible to non-experts

Imagine you're cooking in a kitchen, and you need to choose a method to cook, but you don't know the exact characteristics of the ingredients. This method is like providing a set of different cooking methods, each with its pros and cons. You need to choose the most suitable method based on the possible characteristics of the ingredients. This way, even if you don't know the exact characteristics, you can still make a delicious dish.

ELI14 Explained like you're 14

Imagine you're playing a game, and you need to choose a character to complete a mission, but you don't know the exact characteristics of the enemies. This method is like providing a set of different characters, each with its pros and cons. You need to choose the most suitable character based on the possible characteristics of the enemies. This way, even if you don't know the exact characteristics, you can still successfully complete the mission.

Glossary

Wasserstein Distance

A distance metric used to measure the difference between two probability distributions, often used to quantify design bias.

Used to quantify design bias, ensuring closeness to launch exposure distribution.

Lipschitz Exposure Response

An assumption describing the rate of change in exposure response, ensuring geometric guarantees for design bias.

Used to provide geometric guarantees for design bias.

Exposure Bias

The difference between the exposure distribution caused by a design choice and the target exposure distribution.

A component of planning risk.

Assignment-Unit Variance

Measures the variance of assignment units in an experiment design, affecting the experiment's detection capability.

Used to evaluate the detection capability of experiment designs.

Minimum Detectable Effect

The smallest effect size detectable in an experiment design, affecting the experiment's sensitivity.

Used to evaluate the sensitivity of experiment designs.

Open Questions Unanswered questions from this research

  • 1 How to apply this framework to larger design catalogs?
  • 2 How to better estimate uncertainty sets in practical applications?

Applications

Immediate Applications

Ad System Optimization

Advertisers can use this framework to optimize ad placement strategies under uncertain interference mechanisms, improving ad effectiveness.

Long-term Vision

Recommendation System Improvement

Recommendation systems can use this framework to optimize recommendation algorithms under uncertain interference mechanisms, enhancing user experience.

Abstract

Online experiments in ads, recommendation, and member-experience systems are often planned before the dominant interference mechanism is known. A treatment may propagate through budgets, inventory, producer exposure, graph spillovers, or temporal carryover, making the randomization design itself a statistical decision. We formulate this problem as robust design selection over uncertain exposure mechanisms. Given a finite catalog of six implementable designs, the selector compares each design by worst-case planning risk over an ambiguity set. The risk combines exposure bias, assignment-unit variance, minimum detectable effect, contamination or carryover, operational cost, and estimand mismatch. For theoretical justification, the paper develops a geometry-aware guarantee, stating that design bias is bounded by Wasserstein distance to the launch exposure distribution, and this penalty is minimax tight under Lipschitz exposure response. We also prove finite-catalog approximation and a robust selector theorem with excess-risk control, exact recovery under separation, and certified shortlists when the risk surface is flat. Empirically, the same selector gives different recommendations across samples from public datasets. It selects user-randomization on Criteo ads with dimensionless robust risk 1.295, switchbacks on Open Bandit-bts/men with risk 2.105, and cluster-randomization on KuaiRand with risk 2.240. The Open Bandit case stresses known but uneven logging support, with propensities from 0.00006 to 0.594 and a 5.17% IPS effective-sample share. Overall, the paper contributes an interference-aware experiment design framework based on mechanism-robust design decisions, where the output is either a justified design choice or an uncertainty shortlist.

stat.ML cs.LG