Reliability and Evaluation Research · Jul 2026
causal-inference-engine: Effect Estimation That Can Decline
On the benchmark that matters most, the useful output was no number at all
The problem
Correlation dashboards get read as causal claims.
Approach
Causal discovery classifies every covariate as confounder, mediator or collider and drops the dangerous ones automatically, in plain English, before anything is estimated. Five estimators are compared side by side — propensity score matching, IPW, T/S/X-learner, difference-in-differences and instrumental variables — and graded on synthetic data with a planted true effect, where 10 of 10 rows behaved as expected. Refusal is structural rather than decorative: the interface withholds results until the assumptions page is opened, and the PDF leads with the verdict rather than appending it.
How it works
- The assumptions are drawn as a graph, and every covariate is sorted into confounder, mediator or collider. The dangerous ones are dropped automatically before anything is estimated.
- Feasibility is checked before any effect size: propensity model AUC and the share of units outside common support come first.
- Five estimators run side by side, because agreement across methods relying on different assumptions is stronger evidence than any single number.
- Every estimate carries its estimand. Matching estimates the effect on the treated, weighting across everyone, an instrument only on the units it moved.
- A placebo test replaces the treatment with random noise; any pipeline manufacturing signal fails it.
Key decisions
- Make refusal structural, not decorative
- Caveats get skipped. Instead the interface withholds the effect size until the assumptions page is opened, and the PDF leads with the verdict rather than appending it, because a number seen first is remembered regardless of what follows.
- Grade against planted truth
- On real data nobody knows the true effect, which is why you are estimating it. So the engine is validated on synthetic data with the answer built in: 10 of 10 rows behaved as expected, including one designed to fail.
- Publish the row where the methods miss
- A realistic row with a noisy proxy confounder and another missing entirely lands at 0.83 to 1.07 against a true 0.15, with intervals that confidently exclude the truth. Publishing only the clean rows would misrepresent what these methods deliver.
Interface

What the measurements showed
- On LaLonde, the canonical benchmark, the naive comparison reports −$15,578 against an experimental truth of +$1,794 — wrong by more than $17,000, and in the wrong direction. Matching claws back to −$60. The engine declines to quote an effect at all: propensity AUC 0.970, 47.4% of units outside common support.
- On the no-effect row a raw comparison confidently reports 4.20, while the engine reports 0.03 and marks it non-significant.
- The X-learner recovered per-segment effects at 0.997 correlation with planted truth, including a segment the treatment actively harmed (−1.99 against −2.00) that the +1.9 average conceals.
What broke
- The covariate classifier labelled an instrument as a confounder. The logic asked whether the variable reaches the outcome, but every cause of the treatment reaches the outcome through the treatment. The fix was to ask whether a path exists that avoids the treatment. Small distinction, completely different adjustment set.
- The validation harness graded the instrumental row as a failure, when that row exists precisely to show adjustment has limits. Reporting a correct demonstration as a defect is wrong. Rows now declare their own intent: recover, improve, or fail.
- The placebo test compared the placebo estimate against a fraction of the original, which fails every correct null result since that fraction is near zero by definition. It reported a correctly detected absence of effect as a failure. The fix was to judge against the spread of placebo estimates themselves.
Running it
python -m validation.ground_truthFull setup, configuration and API reference are in the repository README.
Stack
- Python
- DoWhy
- EconML
- CausalML
- statsmodels
- NetworkX
- Streamlit