Skip to content
Back to projects

01 / Context

The problem

Model comparisons are difficult to interpret when training and evaluation pipelines differ between experiments.

02 / Implementation

What I built

Compared BERT, BiDAF, and DistilBERT with a shared training and evaluation workflow across accuracy, precision, recall, and F1.

Current status
Research
Evidence path
Not supplied yet

03 / Current outcome

What exists today

A benchmark-oriented foundation for later applied retrieval work.

Boundary

What this page does not claim

No public deployment or repository link is claimed yet.

Continue exploring

Full archive