Reliability and Evaluation Research · Jul – Aug 2026
personalization-feedback-engine: Online Learning Recommender
Hit rate 0.045 → 0.390 per user, and still last at population level
The problem
Recommenders retrain nightly. What happens if the model updates inside the request that reports the click?
Approach
Per-interaction online learning with River. Weighted signals: purchase 2.0, click 1.0, skip 0.4, impressions excluded. Explanations are validated against real history, so a reason naming an unchosen item is rejected. Every rate carries its standard error, and the experiment layer refuses to return a verdict below 1,000 impressions per arm. 49 tests.
How it works
- The model updates inside the request that reports the interaction, so the next recommendation already includes the click. No batch boundary, no retrain button.
- Signals are weighted: purchase 2.0, click 1.0, skip 0.4, impressions excluded.
- Collaborative similarity is shrunk toward zero when overlap is thin, because a single coincidental co-click is not evidence of anything.
- Every recommendation carries a reason validated against real history, and the system says plainly when it is serving popularity because it does not know the user yet.
- Click-through is attributed back to the exact list that earned it, so per-strategy performance comes from real outcomes.
Key decisions
- Refuse a verdict on thin data
- The experiment layer returns no winner below 1,000 impressions per arm, and every rate carries its standard error. Two strategies a point of CTR apart over a few hundred impressions have not been distinguished.
- Publish the comparison it loses
- At population level the online learner came last: content 0.525, collaborative 0.427, hybrid 0.344, online 0.226. It only wins above roughly 11 events per user. Hiding that would have made the per-user result look like a general one.
- Say what the offline numbers are worth
- They are computed against simulated users with clean preference vectors over 300 items. They establish that the mechanism works. They are not a forecast of live performance, and the README says so before quoting any of them.
Interface

What the measurements showed
- Per-user learning worked: 20 of 20 users improved over 300 interactions, with zero retraining runs.
- Under drift, the online learner recovered +0.082 from feedback alone while batch moved +0.000.
- At population level it came last: content 0.525, collaborative 0.427, hybrid 0.344, online 0.226. It loses at roughly 11 events per user. Published rather than hidden.
What broke
- A collaborative filtering model scored below the random floor because one coincidental co-click created a similarity.
Running it
docker compose up --buildFull setup, configuration and API reference are in the repository README.
Stack
- Python
- River
- Redpanda
- FastAPI
- Next.js
- PostgreSQL
- Redis
- Playwright