Skip to content

Applied ML and Data Science · Mar 2026

Phi-4 Mini Fine-tuned for SEC Financial Q&A

+69% ROUGE-L, trained in three hours on two free-tier T4s

Source

The problem

General-purpose models answer financial filing questions vaguely and confidently.

Approach

QLoRA fine-tuning of Microsoft Phi-4 Mini (3.8B) with 4-bit quantisation on consumer hardware, trained on SEC 10-K financial Q&A. Custom evaluation pipeline measuring semantic similarity and factual accuracy rather than loss alone.

How it works

Key decisions

Evaluate on the task, not the loss curve
Training loss says the model fit the data. It says nothing about whether an answer about a filing is factually right, so the evaluation scores semantic similarity and factual accuracy separately.
Consumer hardware as a constraint, not an excuse
Full fine-tuning a 3.8B model needs 30GB+ of VRAM. QLoRA fits it into two 16GB T4s available free on Kaggle, so the result is reproducible by anyone rather than by anyone with a rented A100.

What the measurements showed

Running it

pip install -r requirements.txt

Full setup, configuration and API reference are in the repository README.

Stack

Written up in