Applied ML and Data Science · Feb – Mar 2026
E-Commerce Product Classifier
Macro F1 0.641 across 19 imbalanced categories, at ~42ms
The problem
Classifiers degrade silently when incoming product data drifts away from the training distribution.
Approach
Fine-tuned DistilBERT across 19 product categories with FastAPI serving. Real-time drift detection via Evidently AI alerts when incoming data diverges from the training distribution, catching silent accuracy degradation before it reaches production. React analytics dashboard, fully containerised.
How it works
- DistilBERT fine-tuned on 50,000 Amazon product samples across 19 categories, served through FastAPI at roughly 42ms per inference.
- Evidently AI tracks text distribution shift using Population Stability Index against the training reference.
- A chi-square framework compares model versions before one is promoted to production.
- Drift alerts fire when the incoming distribution diverges, catching silent accuracy decay before it reaches users.
- A React analytics dashboard renders live classification and drift metrics; the whole stack runs from one Docker command.
Key decisions
- Put drift detection on the serving path
- A classifier that degrades silently is worse than one that fails loudly, because nobody investigates a number that is merely drifting. Checking distribution on the way in makes decay visible while it is still cheap to fix.
- Report the macro F1, not the best category
- Macro F1 of 0.641 across 19 imbalanced categories is the honest aggregate. Digital Music reaches 97.2% and the semantically overlapping categories drag it down, so the per-category breakdown is published rather than the headline figure alone.
What the measurements showed
- Fine-tuned on 50,000 Amazon product samples. Macro F1 0.641 across 19 heavily imbalanced categories, which is the honest aggregate rather than the flattering one.
- Performance splits sharply by category: Digital Music 97.2% F1, Amazon Fashion 85.8%, Automotive 84.1%.
- The weak categories are the semantically overlapping ones, Electronics against Computers and Toys against Baby Products. That is a limitation of flat multi-class taxonomy, and hierarchical classification is the fix.
- Drift is measured with Population Stability Index, and a chi-square framework compares model versions before promotion.
Running it
docker compose upFull setup, configuration and API reference are in the repository README.
Stack
- PyTorch
- Hugging Face
- FastAPI
- Evidently AI
- Docker
- React