Conformal Inference of Counterfactuals and Individual Treatment Effects

TL;DR

Weighted split-CQR delivers finite-sample or doubly robust coverage for counterfactual and ITE intervals.

stat.ME 🔴 Advanced 2020-06-11 25 views
Lihua Lei Emmanuel J. Candès
causal inference conformal inference counterfactuals ITE propensity scores

Key Findings

Methodology

The paper targets potential outcomes Y(1), Y(0), and τ=Y(1)-Y(0). Its core procedure, Weighted split-CQR, fits lower and upper conditional quantiles on a training split, computes calibration scores Vi=max{qlo(Xi)-Yi,Yi-qhi(Xi)}, and calibrates them with propensity-score or density-ratio weights. The output is [qlo(x)-η,qhi(x)+η].

Key Results

  • For completely randomized or stratified randomized trials with perfect compliance, the propensity score is known and the method guarantees average coverage at least 1-α in finite samples, even under a misspecified quantile model; the upper error term is cn^(1/r-1).
  • For observational studies or randomized trials with ignorable compliance, coverage is approximately controlled when either the propensity score or the conditional quantiles are accurately estimated. With weight error Δw, the stated lower bound is 1-α-Δw.
  • Synthetic and real-data experiments report substantial undercoverage from existing CATE intervals, ITE prediction intervals, and Bayesian approaches even in simple models, whereas the proposed intervals attain target coverage with reasonable length. The supplied text does not give dataset names or numerical table values.

Significance

The work shifts causal uncertainty quantification from population averages and CATE point estimates toward individual counterfactuals and treatment effects. This matters when an intervention benefits most people but harms a minority. Distribution-free finite-sample guarantees are more operationally relevant than asymptotic normality when samples are limited, models are misspecified, or covariates are high-dimensional. The framework also connects individualized inference with transportability and target-population decision making.

Technical Contribution

The authors combine Tibshirani et al.'s Weighted Conformal Inference with Romano et al.'s Conformalized Quantile Regression. Propensity scores act as likelihood-ratio corrections between observed treatment groups and target covariate distributions. Proposition 1 gives distribution-free coverage for known weights and quantifies effects of estimated-weight error Δw, moment conditions, and calibration size. A common weighting language covers ATE, ATT, ATC, and generalizability.

Novelty

The key novelty is not merely applying conformal prediction to causal data. It reframes missing-potential-outcome inference as prediction under covariate shift, while covering stratified trials, observational studies, and external target populations. Unlike methods focused mainly on CATE estimation or asymptotic confidence bands, it directly targets marginal coverage for counterfactuals and ITEs and establishes a causal doubly robust property.

Limitations

  • ITE intervals formed from two counterfactual intervals do not identify the joint distribution of Y(1) and Y(0). For individuals outside the study, both outcomes are missing, so the resulting intervals may be conservative.
  • Observational validity still requires strong ignorability and overlap. Propensity scores near zero or one create unstable, very large weights and can produce excessively wide intervals.

Future Work

Important directions include conditional rather than only marginal coverage, dependent data, multi-valued or continuous treatments, sharper joint ITE intervals, and principled weight truncation. Future studies should compare quantile learners and calibration splits systematically, quantify efficiency losses, and extend the framework to causal diagrams and invariant prediction.

AI Executive Summary

Average treatment effects can conceal serious individual risk: a drug may help 70% of patients while worsening the remaining 30%. Yet modern machine-learning causal analyses usually estimate CATE point functions, while asymptotic, bootstrap, or Bayesian intervals can undercover badly in finite samples and under model misspecification.

Lei and Candès propose Weighted split-CQR, combining conformalized quantile regression with covariate-shift correction. Quantile models produce lower and upper outcome predictions; calibration residuals determine an expansion amount η; propensity scores or density ratios reweight calibration observations to the ATE, ATT, ATC, or target-population estimand. Contrasting a missing-outcome interval with the observed outcome yields an ITE interval.

The theory gives finite-sample marginal coverage for randomized trials with known assignment probabilities, and approximate doubly robust coverage in observational settings: either the propensity score or the conditional quantiles may be estimated accurately. Synthetic and real-data experiments show that existing methods exhibit substantial coverage deficits, whereas the proposed intervals reach the target with reasonable length. The supplied paper text does not report dataset names, sample sizes, or table-level numerical percentages, so no unverified values are added.

Deep Analysis

Background

ATE summarizes a population, while CATE, τ(x)=E[Y(1)-Y(0)|X=x], still ignores outcome-level heterogeneity. Regression, normal approximations, resampling, and Bayesian procedures often prioritize point estimation and struggle to provide reliable finite-sample intervals. The potential-outcome formulation adds a fundamental missing-data problem: each unit reveals only one of Y(1) and Y(0).

Core Problem

The targets are C1(x), C0(x), and CITE(x), satisfying P{Y(t)∈Ct(X)}≥1-α and P{τ∈CITE(X)}≥1-α. Obstacles include missing counterfactuals, treatment-induced covariate shift, estimated propensity scores, and the impossibility of nontrivial distribution-free conditional coverage without assumptions.

Innovation

First, Weighted split-CQR treats counterfactual prediction as prediction under covariate shift. Second, inverse propensity weights unify ATE, ATT, ATC, and transportability. Third, randomized trials receive finite-sample distribution-free coverage, while general studies obtain a doubly robust guarantee: accurate propensity scores or accurate conditional quantiles can suffice.

Methodology

  • �� Assume SUTVA, i.i.d. sampling, and strong ignorability; define τ=Y(1)-Y(0).
  • �� Split data into training and calibration folds and fit qαlo and qαhi by quantile regression.
  • �� Compute Vi=max{qαlo(Xi)-Yi,Yi-qαhi(Xi)} on calibration observations.
  • �� Reweight scores using Wi and take a weighted (1-α)-quantile η.
  • �� Return [qαlo(x)-η,qαhi(x)+η].
  • �� Use dQX/dPX together with 1/e(x) or 1/[1-e(x)] for target-population inference.
  • �� Obtain ITE intervals by contrasting counterfactual intervals with observed outcomes, or generalize in-study ITE intervals to new subjects.

Experiments

The paper compares the proposed conformal procedures with existing CATE confidence intervals, ITE prediction intervals, and Bayesian methods on synthetic and real data. It evaluates empirical coverage and interval length under simple smooth settings and model misspecification. The supplied excerpt omits dataset names, sample sizes, α values, hyperparameters, and numerical baseline tables, so those details cannot be reconstructed without inventing evidence.

Results

With known weights, Proposition 1 guarantees coverage at least 1-α; under no ties and finite r-th weight moments, the upper bound is 1-α+cn^(1/r-1). Estimated weights introduce Δw, giving a lower bound 1-α-Δw. Empirically, baseline procedures substantially undercover, whereas the conformal method reaches nominal coverage with reasonably short intervals; the supplied material contains no dataset-specific numbers.

Applications

Clinical decision systems can report patient-level ranges for untreated or treated outcomes rather than only average efficacy. Policy and education studies can target ATT, ATC, or an external population. Prerequisites are measured confounders, adequate overlap, and credible propensity-score or conditional-quantile estimation; outputs should include both coverage diagnostics and interval length.

Limitations & Outlook

Strong ignorability is not testable from observed data, and weak overlap causes exploding weights and unstable calibration. Marginal coverage does not ensure conditional validity for every covariate profile. ITE joint uncertainty remains unidentified because Y(1) and Y(0) are never jointly observed. Data splitting can reduce efficiency, while poor quantile or propensity models increase length. Future work should target conditional coverage, joint calibration, complex treatments, and robust weighting.

Plain Language Accessible to non-experts

Imagine a school testing two study programs. A student can attend tutoring or regular class, but cannot experience both at the same time. We want to know not only the average score change, but the range of scores this particular student might have achieved under the program they did not receive.

The method first studies students whose answers are known and predicts a reasonable score range rather than one exact score. It then checks the prediction on another group and measures how far the real scores fall outside the ranges. If some kinds of students were rarely assigned to tutoring, their checking results receive more weight, so the final range better represents the students we care about.

The useful safety feature is that the method can still work when either the assignment probabilities are estimated well or the score ranges are estimated well. It does not require both to be perfect. However, if a type of student almost never appears in one program, the range must become very wide. And because no student can live both possible school histories, the method cannot reveal the exact connection between the two unseen scores.

ELI14 Explained like you're 14

Think of testing two game weapons: a fire sword and an ice sword. Each player can try only one, so for every player one possible result is missing. Saying “the fire sword scores 20 more damage on average” may hide the fact that it helps some players and hurts others.

This paper builds a realistic score range for each player. It learns from players who already used a weapon, then checks its guesses on another group. If the guesses miss by a lot, the range expands. It also corrects for the fact that some types of players were much more likely to receive one weapon than the other.

Here is the clever part: the whole system can still be trustworthy if either the weapon-assignment chances are estimated well or the score ranges are estimated well. That is the paper’s doubly robust idea. We do not need every model to be perfect!

But it is not magic armor. If almost nobody like you used the other weapon, the prediction becomes uncertain and the range gets wide. Also, one player cannot play both alternate histories at once, so the exact personal effect can never be directly observed. The method gives a careful range, not a guaranteed prophecy.

Glossary

Potential outcome

The two outcomes a unit could have under treatment and control, Y(1) and Y(0). Only one is observed for each unit.

The paper defines counterfactuals and ITEs through this framework.

Individual Treatment Effect (ITE)

τ=Y(1)-Y(0), the treatment-control contrast for one individual. It is generally unobservable because the two potential outcomes are never jointly observed.

The principal target of interval inference.

Weighted split-CQR

A split conformal method that fits lower and upper conditional quantiles and calibrates their errors using covariate weights. It adapts prediction intervals to a target distribution.

The paper's central algorithm.

Propensity score

e(x)=P(T=1|X=x), the probability of treatment conditional on covariates. Inverse propensity weights correct treatment-selection differences.

It determines ATE, ATT, ATC, and transportability weights.

Strong ignorability

Conditional on X, treatment assignment is independent of both potential outcomes, together with an appropriate overlap condition. It excludes unmeasured confounding.

The main observational-study assumption.

Marginal coverage

The probability, over a randomly drawn target unit, that the interval contains the outcome is at least 1-α. It does not imply validity at every fixed covariate value.

The formal coverage criterion throughout the paper.

Open Questions Unanswered questions from this research

  • 1 How can one obtain useful conditional coverage for specific subgroups without making intervals prohibitively wide? The paper emphasizes marginal guarantees because distribution-free conditional guarantees are generally impossible without additional structure.
  • 2 How should the method behave under unmeasured confounding, near-violations of overlap, continuous treatments, or strongly dependent observations? These settings require sensitivity analysis, stabilized weighting, and new conformal constructions.

Applications

Immediate Applications

Personalized clinical treatment

Hospitals can estimate a patient's missing-treatment outcome interval and contrast it with the observed or predicted alternative. Deployment requires measured confounders, overlap, reliable propensity or quantile models, and reporting of empirical coverage and interval width.

Policy and education evaluation

Agencies and schools can target ATT, ATC, or an external population and report ranges for people who did not receive an intervention. This supports identifying likely beneficiaries, neutral cases, and possible harms beyond an average effect.

Long-term Vision

Auditable individualized decision systems

Conformal intervals could accompany recommendations in healthcare and public services, exposing uncertainty rather than hiding it. Major obstacles include unmeasured confounding, privacy, distribution shift, computational monitoring, and responsibility for decisions.

Abstract

Evaluating treatment effect heterogeneity widely informs treatment decision making. At the moment, much emphasis is placed on the estimation of the conditional average treatment effect via flexible machine learning algorithms. While these methods enjoy some theoretical appeal in terms of consistency and convergence rates, they generally perform poorly in terms of uncertainty quantification. This is troubling since assessing risk is crucial for reliable decision-making in sensitive and uncertain environments. In this work, we propose a conformal inference-based approach that can produce reliable interval estimates for counterfactuals and individual treatment effects under the potential outcome framework. For completely randomized or stratified randomized experiments with perfect compliance, the intervals have guaranteed average coverage in finite samples regardless of the unknown data generating mechanism. For randomized experiments with ignorable compliance and general observational studies obeying the strong ignorability assumption, the intervals satisfy a doubly robust property which states the following: the average coverage is approximately controlled if either the propensity score or the conditional quantiles of potential outcomes can be estimated accurately. Numerical studies on both synthetic and real datasets empirically demonstrate that existing methods suffer from a significant coverage deficit even in simple models. In contrast, our methods achieve the desired coverage with reasonably short intervals.

stat.ME math.ST stat.ML