Skip to content

Reliability and Evaluation Research · Jul – Aug 2026

personalization-feedback-engine: Online Learning Recommender

Hit rate 0.045 → 0.390 per user, and still last at population level

Source

The problem

Recommenders retrain nightly. What happens if the model updates inside the request that reports the click?

Approach

Per-interaction online learning with River. Weighted signals: purchase 2.0, click 1.0, skip 0.4, impressions excluded. Explanations are validated against real history, so a reason naming an unchosen item is rejected. Every rate carries its standard error, and the experiment layer refuses to return a verdict below 1,000 impressions per arm. 49 tests.

How it works

Key decisions

Refuse a verdict on thin data
The experiment layer returns no winner below 1,000 impressions per arm, and every rate carries its standard error. Two strategies a point of CTR apart over a few hundred impressions have not been distinguished.
Publish the comparison it loses
At population level the online learner came last: content 0.525, collaborative 0.427, hybrid 0.344, online 0.226. It only wins above roughly 11 events per user. Hiding that would have made the per-user result look like a general one.
Say what the offline numbers are worth
They are computed against simulated users with clean preference vectors over 300 items. They establish that the mechanism works. They are not a forecast of live performance, and the README says so before quoting any of them.

Interface

Side-by-side comparison of the four recommendation strategies
Each strategy measured under identical conditions, including the population-level comparison the online learner loses.

What the measurements showed

What broke

Running it

docker compose up --build

Full setup, configuration and API reference are in the repository README.

Stack

Written up in