Skip to content

Repository files navigation

False Discovery Rate Lab

A reproducible Monte Carlo audit of multiple-testing procedures when hypotheses are correlated in blocks. The project separates three quantities that are often blurred together in exploratory research: false discovery rate (FDR), family-wise error rate (FWER), and power.

Five decision rules are compared at a nominal 5% level:

  • Uncorrected: reject every raw p-value below 0.05.
  • Bonferroni: single-step FWER control.
  • Holm: step-down FWER control.
  • Benjamini-Hochberg (BH): step-up FDR control under independence and common forms of positive dependence.
  • Benjamini-Yekutieli (BY): conservative FDR control under arbitrary dependence.

Simulation design

Each replication contains 200 two-sided Gaussian z-tests arranged in ten blocks of 20. Tests share a Gaussian shock within their block, giving correlations of 0, 0.3, or 0.7. Two scenarios are evaluated:

  1. Global null: all 200 hypotheses are null, used to audit false positives.
  2. Sparse signals: 20 non-null hypotheses (two per block) have a z-scale mean shift of 3, used to estimate FDR and power.

The checked-in run uses 5,000 replications for every scenario-correlation pair and a fixed seed. All data are deterministic synthetic draws; no external or private data are used.

Reproduce

python -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[test]"
pytest -W error
python -m false_discovery_rate_lab --output-dir artifacts

Results

At within-block correlation 0.3, the held-out simulation estimates were:

Method Sparse-signal FDR Sparse-signal power Global-null FWER
Uncorrected 0.3339 0.8484 0.9996
Bonferroni 0.0073 0.2546 0.0474
Holm 0.0074 0.2567 0.0474
BH 0.0450 0.4979 0.0514
BY 0.0066 0.2518 0.0080

Across correlations 0, 0.3, and 0.7, BH's empirical FDR ranged from 0.0431 to 0.0451 and power ranged from 0.4975 to 0.4989. In this design it roughly doubled the power of the FWER-controlling procedures while remaining below the 5% FDR target. The uncorrected rule reached about 85% power, but 30.9%–33.9% of its discoveries were false on average.

Under the global null, BH's largest observed FWER was 0.0514 (Monte Carlo standard error 0.0031), so the small excess over 0.05 is not distinguishable from simulation noise at this resolution. BY's harmonic correction made it nearly as conservative as Bonferroni here. These findings describe the stated positive block-dependence design; they are not evidence that BH controls FDR under every dependence structure.

Empirical FDR across dependence levels

Outputs

  • artifacts/simulation_results.csv: FDR, FWER, power, Monte Carlo standard errors, and mean rejection counts for all 30 design cells.
  • artifacts/run_summary.json: machine-readable design and headline metrics.
  • artifacts/fdr_vs_dependence.png: false-discovery behavior with sparse signals.
  • artifacts/power_vs_dependence.png: power across dependence levels.
  • artifacts/all_null_fwer.png: chance of at least one false rejection under the global null.

Repository structure

src/false_discovery_rate_lab/
  procedures.py   # vectorized Bonferroni, Holm, BH, and BY rules
  simulation.py   # block-correlated Gaussian tests
  evaluation.py   # Monte Carlo metrics and orchestration
  reporting.py    # static figures
tests/             # procedure, simulation, metric, and E2E tests
artifacts/         # reproducible outputs from the default run

Limitations

  • Gaussian block dependence is a controlled stress test, not a model of every scientific or business testing pipeline.
  • The effect size and 10% signal prevalence are fixed; power is not universal.
  • BH has theoretical guarantees for independence and certain positive dependence structures, while BY is designed for arbitrary dependence. This experiment is an empirical check of selected settings, not a replacement for those conditions.
  • Monte Carlo estimates have sampling error; the CSV reports standard errors so small differences are not over-interpreted.
  • Selection-adjusted effect estimation and hierarchical testing are out of scope.

Responsible use

This is an educational portfolio project. Statistical thresholds should be chosen from the consequences and design of a real analysis, not copied from this simulation.

About

Monte Carlo audit of multiple-testing procedures under block dependence

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages