Applied ML and Data Science · Jan 2026
Hybrid Recommendation Engine
129,782 ratings at 91.35% sparsity, three signals blended
The problem
Collaborative filtering has nothing to say about a user or item it has never seen.
Approach
Collaborative filtering, content-based filtering and SVD matrix factorisation blended at a fixed 0.4/0.3/0.3 to address the cold start problem, and evaluated on catalog coverage, category diversity and average rating rather than raw accuracy.
How it works
- Collaborative filtering, content-based filtering and SVD matrix factorisation each produce candidate recommendations.
- A hybrid layer blends them at a fixed 0.4 collaborative, 0.3 content-based, 0.3 SVD, which is what covers the cold start neither pure approach solves alone.
- Quality is judged on catalog coverage, category diversity and the average rating of what gets recommended, rather than on raw accuracy.
- A Streamlit dashboard lets you inspect recommendations per user.
Key decisions
- Hybrid specifically for cold start
- Collaborative filtering knows nothing about a new item and content-based filtering only ever suggests more of the same. Combining them covers each other's blind spot rather than averaging two mediocre answers.
- Judge the list, not the prediction
- Accuracy on a held-out rating says little about a short ranked list someone actually sees. Coverage, diversity and the average rating of served recommendations correspond to the product surface instead.
What the measurements showed
- 3,000 users, 500 products across 7 categories, 129,782 ratings, 91.35% sparsity.
- Hybrid score is a fixed blend: 0.4 collaborative, 0.3 content-based, 0.3 SVD.
- Recommendations average above 4.0 of 5.0 on rating, with coverage spread across categories rather than concentrated in popular items.
Running it
streamlit run app.pyFull setup, configuration and API reference are in the repository README.
Stack
- Python
- scikit-learn
- pandas
- NumPy
- SVD