Test Martingales, Bayes Factors and $p$-Values
This paper links test martingales with Bayes factors and p-values, introducing functions to limit evidence exaggeration, enabling systematic conversion between them.
Key Findings
Methodology
The paper systematically analyzes the relationship between nonnegative martingales, Bayes factors, and p-values. It introduces a class of monotone functions f that constrain the maximum of a martingale, ensuring the evidence is not exaggerated. Using martingale convergence theorems and properties of supremum processes, the authors establish conditions under which the inverse of the supremum serves as a valid p-value or Bayes factor. The methodology involves constructing calibration functions satisfying integral conditions, leveraging Fatou’s lemma and Doob’s convergence theorem, to formalize the duality between dynamic evidence measures and static hypothesis metrics.
Key Results
- Proved that the inverse of the limiting value of a test supermartingale equals a Bayes factor, and that the inverse of its supremum can serve as a p-value, with the calibration functions satisfying specific integral constraints.
- Provided a complete characterization of all increasing functions that act as calibrators, ensuring the conversion between p-values and Bayes factors is both valid and conservative.
- Numerical simulations demonstrate that the proposed supremum limiting functions effectively control evidence exaggeration, outperforming traditional methods especially in continuous monitoring scenarios.
Significance
This work advances the theoretical understanding of dynamic evidence measures, bridging the gap between sequential analysis and static hypothesis testing. It offers a unified framework to control evidence exaggeration in multiple testing, with implications for online learning, financial risk management, and clinical trials. The results facilitate more reliable inference in settings where repeated or continuous testing occurs, addressing longstanding issues of p-value misinterpretation and Bayesian-frequentist reconciliation.
Technical Contribution
The paper introduces a novel class of calibration functions for martingale supremum processes, rigorously characterizing their properties and establishing a duality with p-values and Bayes factors. It extends classical martingale convergence results to include supremum-based evidence measures, providing explicit integral conditions for calibration. The work also develops a comprehensive theory for the optimal design of these functions, ensuring minimal evidence exaggeration while maintaining statistical validity, thus opening new avenues for dynamic hypothesis testing.
Novelty
This is the first systematic characterization of all monotone functions serving as calibrators for martingale supremum processes, linking the dynamic and static measures of evidence in a unified theory. Unlike prior work focusing solely on martingale convergence, this study emphasizes the control of extremal behavior, providing a rigorous foundation for evidence calibration in sequential and continuous testing frameworks. It bridges the gap between Bayesian and frequentist approaches through explicit functional transformations.
Limitations
- The theoretical framework assumes the non-negativity and integrability of martingales, which may not hold in certain high-dimensional or non-standard models.
- Designing optimal calibration functions in high-frequency or large-scale applications could be computationally intensive, requiring further algorithmic development.
- Real-world data may introduce noise and deviations from ideal assumptions, potentially affecting the effectiveness of the proposed evidence control strategies, necessitating robust extensions.
Future Work
Future research will focus on extending these calibration techniques to high-dimensional and nonparametric models, integrating machine learning algorithms for adaptive calibration in online settings. Additionally, exploring the application of these methods in complex data environments such as time series, network data, and large-scale experiments will be prioritized. Developing computationally efficient algorithms for real-time evidence calibration remains an important direction.
AI Executive Summary
This study explores the deep connection between test martingales, Bayes factors, and p-values, fundamental tools in statistical hypothesis testing. By analyzing the supremum process of nonnegative martingales, the authors introduce a class of monotone functions—calibrators—that effectively limit the exaggeration of evidence often encountered in sequential testing.
The core idea is to leverage the properties of martingale convergence and supremum processes to establish a duality: the inverse of the limiting value of a supermartingale corresponds to a Bayes factor, while the inverse of its supremum can serve as a p-value. The authors rigorously characterize all functions satisfying the integral conditions necessary for these transformations, ensuring statistical validity and conservativeness.
Numerical experiments demonstrate that these calibration functions significantly reduce evidence exaggeration, especially in continuous monitoring scenarios, outperforming traditional p-value adjustments. This work bridges the gap between dynamic evidence measures and static hypothesis metrics, providing a unified framework that enhances the reliability of sequential hypothesis testing.
The implications are broad: in fields like clinical trials, finance, and machine learning, where ongoing data collection is common, these methods can improve decision-making robustness. The paper also discusses future directions, including extending the framework to high-dimensional models and developing efficient algorithms for real-time applications, promising a substantial impact on both theoretical and applied statistics.
Deep Dive
Abstract
A nonnegative martingale with initial value equal to one measures evidence against a probabilistic hypothesis. The inverse of its value at some stopping time can be interpreted as a Bayes factor. If we exaggerate the evidence by considering the largest value attained so far by such a martingale, the exaggeration will be limited, and there are systematic ways to eliminate it. The inverse of the exaggerated value at some stopping time can be interpreted as a $p$-value. We give a simple characterization of all increasing functions that eliminate the exaggeration.