Applied ML and Data Science · Jan 2026
A/B Testing and Experimentation Framework
A worked decision: +23.06% lift, 95% CI [1.72pp, 3.56pp]
The problem
Teams call experiments early and read noise as signal.
Approach
Sample size estimation, posterior probability and expected impact under uncertainty, so a result reports what it can and cannot support rather than a bare p-value.
How it works
- Sample size is estimated before the experiment runs, so the test is powered for the effect it hopes to detect.
- Frequentist and Bayesian analyses run side by side on the same data.
- Bayesian output reports posterior probability and expected impact under uncertainty rather than a bare significance verdict.
- A translation layer turns the statistics into the decision they support, ending in ship or do not ship rather than a p-value.
Key decisions
- State plainly what the framework does not do
- The README carries an explicit 'does NOT do' section. A tool that appears to answer questions it cannot answer is more dangerous than one with obvious gaps.
- Both paradigms, side by side
- Frequentist significance answers whether to reject a null; posterior probability and expected impact answer whether to ship. Showing both stops the p-value being read as the business decision.
What the measurements showed
- Worked example: control 11.45% against treatment 14.09% on 10,000 users each. Absolute lift +2.64pp, relative +23.06%, p < 0.000001, 95% CI [1.72pp, 3.56pp], Bayesian P(treatment better) approximately 100%.
- The README states what the framework will not do: no sequential testing or early stopping, no automatic correction for multiple simultaneous tests, no non-binary outcome metrics.
Running it
streamlit run app.pyFull setup, configuration and API reference are in the repository README.
Stack
- Python
- SciPy
- NumPy
- pandas