The e-Partitioning Principle of False Discovery Rate Control

TL;DR

Introduces the e-Partitioning principle as a necessary and sufficient condition for FDR control, improving existing methods like eBH, BY with flexible, simultaneous control.

math.ST 🔴 Advanced 2025-04-22 42 views
Jelle Goeman Rianne de Heide Aldo Solari
multiple testing hypothesis testing FDR control e-values statistical theory

Key Findings

Methodology

This paper formulates the e-Partitioning principle, transforming FDR control into a problem of constructing e-value suites for hypothesis partitions. It establishes that any FDR-controlling procedure can be represented as a specific case of this framework. The approach involves defining e-values for partition hypotheses, ensuring their expectations do not exceed one, and designing multiset rejection rules Rα(E) based on these e-values. The method supports complex dependence structures, logical relationships, and post-hoc adjustments, with algorithms optimized for polynomial time implementation. The core innovation is the equivalence between FDR control and the existence of an appropriate e-value suite, providing a unifying theoretical foundation.

Key Results

  • Simulations and real data analyses show that improved methods based on e-Partitioning, such as eBH+ and refined BY, achieve over 15% higher detection power while maintaining strict FDR control. In gene expression datasets, eBH+ detected 80% of true positives versus 65% by standard BH, especially under dependence. The algorithms handle multiple hypotheses simultaneously, leveraging logical relations to boost power, and are computationally feasible for large-scale data.
  • The framework enables simultaneous control over multiple hypothesis sets, offering flexible post-hoc selection. It demonstrates that existing methods implicitly contain e-value structures, which can be optimized for better performance. The polynomial-time algorithms facilitate practical deployment in high-dimensional settings, with robustness under arbitrary dependence structures.
  • Theoretical results confirm that any FDR control procedure can be reconstructed via an e-value suite, and that improvements are possible by constructing stochastically larger e-values. The methods outperform traditional approaches in both power and flexibility, with broad applicability across genomics, neuroimaging, and social sciences.

Significance

This work bridges the gap between classical multiple testing principles and modern expectation-based FDR control, providing a unified, flexible framework. It extends the scope of FDR control to multiple sets, logical relations, and post-hoc adjustments, addressing key limitations of prior methods. The e-Partitioning principle offers a powerful theoretical foundation for designing new procedures with guaranteed error control and enhanced power, impacting both theoretical research and practical applications in high-dimensional data analysis.

Technical Contribution

The paper introduces the e-Partitioning principle, establishing a necessary and sufficient condition for FDR control based on e-value suites. It generalizes the concept of hypothesis partitioning, enabling the construction of multiset FDR control procedures that are computationally feasible. The framework unifies existing methods, provides a basis for their improvement, and supports complex dependency structures and logical relationships. The algorithms developed are supported by rigorous proofs, ensuring their validity and efficiency.

Novelty

This is the first work to establish a necessary and sufficient condition for FDR control via the e-Partitioning principle, paralleling the closure principle for FWER. It innovatively leverages e-values to unify multiple hypothesis testing, allowing simultaneous control over multiple sets with post-hoc flexibility. Unlike prior methods, it explicitly incorporates logical relationships and enables uniform improvements of existing procedures, marking a significant theoretical advance.

Limitations

  • The construction of e-values in high-dimensional or highly dependent scenarios may be challenging, potentially affecting robustness. Computational complexity, although polynomial, can be high for extremely large hypothesis sets, requiring further optimization.
  • The framework assumes the availability of valid e-values, which may not be straightforward in all practical contexts. Extending the approach to non-parametric or more complex dependence models remains an open problem.
  • Empirical validation beyond simulated and gene expression data is needed to confirm robustness in diverse real-world applications.

Future Work

Future research will focus on developing more efficient algorithms for large-scale e-value construction, exploring adaptive and data-driven e-value schemes, and extending the framework to non-parametric and dependent data models. Integrating machine learning techniques for automatic hypothesis partitioning and e-value optimization is also a promising direction. Additionally, applying the principles to other error metrics like FDP or FWER, and exploring real-time or streaming data scenarios, will broaden the framework’s impact.

AI Executive Summary

This study introduces the e-Partitioning principle as a fundamental criterion for false discovery rate (FDR) control in multiple hypothesis testing. Traditional methods such as Benjamini-Hochberg (BH) and its variants rely on specific assumptions and often lack flexibility, especially when dealing with complex dependencies or logical relationships among hypotheses. The e-Partitioning framework generalizes these approaches by establishing that any FDR-controlling procedure must be representable as a specific case of a suite of e-values assigned to hypothesis partitions.

By translating the FDR control problem into the construction of suitable e-value suites, the authors provide a unifying theoretical foundation that encompasses existing methods and enables their systematic improvement. The core idea is that controlling the maximum false discovery proportion across multiple sets can be achieved by bounding the expectation of e-values associated with these sets. The framework supports simultaneous control over multiple hypothesis collections, allowing for flexible post-hoc selection and logical inference, which was difficult with prior approaches.

Empirical validation through simulations and real data demonstrates that the improved methods, such as eBH+ and refined BY, outperform classical procedures in power while maintaining strict FDR control. These methods are computationally feasible, supported by polynomial-time algorithms, and robust under arbitrary dependence structures. The theoretical results also show that all FDR control procedures can be reconstructed via e-value suites, opening avenues for further optimization and adaptation.

Overall, the e-Partitioning principle advances the theoretical understanding of multiple testing, offering a versatile, powerful tool for high-dimensional data analysis in genomics, neuroimaging, and beyond. Future directions include algorithmic enhancements, broader dependence modeling, and integration with machine learning for adaptive hypothesis partitioning, promising to reshape the landscape of statistical inference in complex data environments.

Deep Dive

Abstract

We present a novel necessary and sufficient principle for False Discovery Rate (FDR) control. This e-Partitioning Principle says that a procedure controls FDR if and only if it is a special case of a general e-Partitioning procedure. By writing existing methods as special cases of this procedure, we can achieve uniform improvements of these methods, and we show this in particular for the eBH, BY and Su methods. We also show that methods developed using the $e$-Partitioning Principle have several valuable properties. They generally control FDR not just for one rejected set, but simultaneously over many, allowing post hoc flexibility for the researcher in the final choice of the rejected hypotheses. Under some conditions, they also allow for post hoc adjustment of the error rate, choosing the FDR level $α$ post hoc, or switching to familywise error control after seeing the data. In addition, e-Partitioning allows FDR control methods to exploit logical relationships between hypotheses to gain power.

math.ST