Applied ML and Data Science · Mar 2026
Phi-4 Mini Fine-tuned for SEC Financial Q&A
+69% ROUGE-L, trained in three hours on two free-tier T4s
The problem
General-purpose models answer financial filing questions vaguely and confidently.
Approach
QLoRA fine-tuning of Microsoft Phi-4 Mini (3.8B) with 4-bit quantisation on consumer hardware, trained on SEC 10-K financial Q&A. Custom evaluation pipeline measuring semantic similarity and factual accuracy rather than loss alone.
How it works
- Microsoft Phi-4 Mini (3.8B) fine-tuned with QLoRA and 4-bit NF4 quantisation on two Tesla T4s from Kaggle's free tier, in about three hours over three epochs.
- LoRA adapters at r=16, alpha=32, dropout 0.05, training roughly 40M parameters against the model's 3.8B.
- Trained on 6,300 SEC 10-K question-answer pairs, where general-purpose models answer vaguely and confidently.
- A custom evaluation pipeline measures semantic similarity and factual accuracy rather than loss alone.
- Reported as a +69% ROUGE-L improvement over the base model on the same benchmark.
Key decisions
- Evaluate on the task, not the loss curve
- Training loss says the model fit the data. It says nothing about whether an answer about a filing is factually right, so the evaluation scores semantic similarity and factual accuracy separately.
- Consumer hardware as a constraint, not an excuse
- Full fine-tuning a 3.8B model needs 30GB+ of VRAM. QLoRA fits it into two 16GB T4s available free on Kaggle, so the result is reproducible by anyone rather than by anyone with a rented A100.
What the measurements showed
- ROUGE-1 0.4657 to 0.7523 (+61.6%), ROUGE-2 0.3560 to 0.6106 (+71.5%), ROUGE-L 0.4242 to 0.7168 (+69.0%), on held-out SEC 10-K pairs.
- Roughly 40M trainable parameters against 3.8B, about 1% of the model, with a final training loss of 1.07 over three epochs.
- Full fine-tuning would have needed 30GB+ of VRAM. QLoRA brought it inside two 16GB T4s on Kaggle's free tier, which makes the result reproducible by anyone.
Running it
pip install -r requirements.txtFull setup, configuration and API reference are in the repository README.
Stack
- Python
- PyTorch
- PEFT/LoRA
- QLoRA
- Hugging Face