A reproducible Monte Carlo audit of multiple-testing procedures when hypotheses are correlated in blocks. The project separates three quantities that are often blurred together in exploratory research: false discovery rate (FDR), family-wise error rate (FWER), and power.
Five decision rules are compared at a nominal 5% level:
- Uncorrected: reject every raw p-value below 0.05.
- Bonferroni: single-step FWER control.
- Holm: step-down FWER control.
- Benjamini-Hochberg (BH): step-up FDR control under independence and common forms of positive dependence.
- Benjamini-Yekutieli (BY): conservative FDR control under arbitrary dependence.
Each replication contains 200 two-sided Gaussian z-tests arranged in ten blocks of 20. Tests share a Gaussian shock within their block, giving correlations of 0, 0.3, or 0.7. Two scenarios are evaluated:
- Global null: all 200 hypotheses are null, used to audit false positives.
- Sparse signals: 20 non-null hypotheses (two per block) have a z-scale mean shift of 3, used to estimate FDR and power.
The checked-in run uses 5,000 replications for every scenario-correlation pair and a fixed seed. All data are deterministic synthetic draws; no external or private data are used.
python -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[test]"
pytest -W error
python -m false_discovery_rate_lab --output-dir artifactsAt within-block correlation 0.3, the held-out simulation estimates were:
| Method | Sparse-signal FDR | Sparse-signal power | Global-null FWER |
|---|---|---|---|
| Uncorrected | 0.3339 | 0.8484 | 0.9996 |
| Bonferroni | 0.0073 | 0.2546 | 0.0474 |
| Holm | 0.0074 | 0.2567 | 0.0474 |
| BH | 0.0450 | 0.4979 | 0.0514 |
| BY | 0.0066 | 0.2518 | 0.0080 |
Across correlations 0, 0.3, and 0.7, BH's empirical FDR ranged from 0.0431 to 0.0451 and power ranged from 0.4975 to 0.4989. In this design it roughly doubled the power of the FWER-controlling procedures while remaining below the 5% FDR target. The uncorrected rule reached about 85% power, but 30.9%–33.9% of its discoveries were false on average.
Under the global null, BH's largest observed FWER was 0.0514 (Monte Carlo standard error 0.0031), so the small excess over 0.05 is not distinguishable from simulation noise at this resolution. BY's harmonic correction made it nearly as conservative as Bonferroni here. These findings describe the stated positive block-dependence design; they are not evidence that BH controls FDR under every dependence structure.
artifacts/simulation_results.csv: FDR, FWER, power, Monte Carlo standard errors, and mean rejection counts for all 30 design cells.artifacts/run_summary.json: machine-readable design and headline metrics.artifacts/fdr_vs_dependence.png: false-discovery behavior with sparse signals.artifacts/power_vs_dependence.png: power across dependence levels.artifacts/all_null_fwer.png: chance of at least one false rejection under the global null.
src/false_discovery_rate_lab/
procedures.py # vectorized Bonferroni, Holm, BH, and BY rules
simulation.py # block-correlated Gaussian tests
evaluation.py # Monte Carlo metrics and orchestration
reporting.py # static figures
tests/ # procedure, simulation, metric, and E2E tests
artifacts/ # reproducible outputs from the default run
- Gaussian block dependence is a controlled stress test, not a model of every scientific or business testing pipeline.
- The effect size and 10% signal prevalence are fixed; power is not universal.
- BH has theoretical guarantees for independence and certain positive dependence structures, while BY is designed for arbitrary dependence. This experiment is an empirical check of selected settings, not a replacement for those conditions.
- Monte Carlo estimates have sampling error; the CSV reports standard errors so small differences are not over-interpreted.
- Selection-adjusted effect estimation and hierarchical testing are out of scope.
This is an educational portfolio project. Statistical thresholds should be chosen from the consequences and design of a real analysis, not copied from this simulation.
