Skip to content

Applied ML and Data Science · Jan 2026

A/B Testing and Experimentation Framework

A worked decision: +23.06% lift, 95% CI [1.72pp, 3.56pp]

Source

The problem

Teams call experiments early and read noise as signal.

Approach

Sample size estimation, posterior probability and expected impact under uncertainty, so a result reports what it can and cannot support rather than a bare p-value.

How it works

Key decisions

State plainly what the framework does not do
The README carries an explicit 'does NOT do' section. A tool that appears to answer questions it cannot answer is more dangerous than one with obvious gaps.
Both paradigms, side by side
Frequentist significance answers whether to reject a null; posterior probability and expected impact answer whether to ship. Showing both stops the p-value being read as the business decision.

What the measurements showed

Running it

streamlit run app.py

Full setup, configuration and API reference are in the repository README.

Stack