Skip to content

Stage 3 · Object detection

Building the YOLO dataset from MySQL and fine-tuning a fish detector, tracked in MLflow.

Govinda Lienart · AquaMind · Stage 3

prepare_dataset.py · YOLOv8 · MLflow Model Registry

Abstract

What this stage produces in one paragraph: a versioned YOLO dataset pulled from MySQL, the fine-tuned detector trained on it, and its registration in the MLflow Model Registry — headline numbers (mAP, dataset size, train/val split) once measured.

Why a dataset needs to be built from MySQL rather than used as flat files (traceability, reproducibility, the ability to re-run prepare_dataset.py against a different set of annotation_set_ids without re-touching raw frames). Where this sits in the pipeline: consumes Stage 2’s annotations table, produces the detector Stage 4 depends on.

State the two things this stage sets out to do: (1) turn annotated MySQL rows into a YOLO-format dataset folder with a reproducible train/val split, (2) fine-tune YOLOv8 on that dataset and register the result so later stages resolve it by name/alias instead of a hardcoded path.

Annotated frames leave MySQL as a YOLO dataset folder, a pretrained YOLOv8 is fine-tuned on that folder in a GPU notebook, and the resulting weights are registered in the MLflow Model Registry, so later stages load the detector by name instead of by file path. The figure below shows the full pipeline.

Stage 3 pipeline. The labelled dataset from Stage 2 (annotations, annotation_sets and frames tables in MySQL) is turned into a YOLO dataset folder by prepare_dataset.py, which picks the annotated frames of the chosen annotation sets, splits them 80/20 into train and validation with a fixed seed, and writes the images and one YOLO label file per frame, plus dataset_card.yaml (annotation sets, videos, counts, git commit) and dataset.yaml. The dataset is versioned with DVC, with the pointer file in git and the data on Google Drive. A Kaggle GPU notebook clones the repository, pulls the dataset with DVC and fine-tunes a pretrained YOLOv8. The run folder (best.pt, last.pt, results.csv, plots) is downloaded by hand. log_artifact_mlflow.py logs the parameters, per-epoch metrics and dataset card to MLflow and registers best.pt as a new version of aquamind-yolo-detector in the MLflow Model Registry, moving the champion alias to it. Later stages load the detector by alias: the tracker in Stage 4 and pre-labelling in Stage 5.
FigStage 3 pipeline: the labelled dataset from Stage 2 becomes a versioned YOLO dataset folder (prepare_dataset.py, DVC), a pretrained YOLOv8 is fine-tuned in a Kaggle GPU notebook, and log_artifact_mlflow.py logs the run and registers best.pt in the MLflow Model Registry, where later stages load it by the @champion alias. The training run is downloaded from Kaggle by hand. Click to zoom.

The first three subsections below follow the pipeline from top to bottom.

Dataset construction — prepare_dataset.py

Section titled “Dataset construction — prepare_dataset.py”

Query annotated frames from MySQL, 80/20 train/val split, folder layout written to dataset/{name}/, dataset_card.yaml (counts, annotation_set_ids, git commit for reproducibility), dataset.yaml for YOLO.

Base model, training config (epochs, image size, augmentation if any), hardware used, what changed between iterations if there were several dataset/training rounds.

MLflow tracking and Model Registry — log_artifact_mlflow.py

Section titled “MLflow tracking and Model Registry — log_artifact_mlflow.py”

Logging the run to MLflow, registering best.pt under aquamind-yolo-detector, moving the @champion alias. Why registry over a hardcoded config.yaml path (promotion = re-run the script). Two-backend note: SQLite registry vs file-based mlruns/ for other experiments.

0 = danio_rerio, 1 = reflection — why reflections are a separate class rather than filtered out (helps the model learn to ignore them rather than never seeing them).

(Once trained, this stage’s @champion model is also used in Stage 5 to pre-label new hard cases for correction, and the relabelled data comes back here for the next training round.)

Final dataset size (train/val counts), class balance, any frame-source mix (regular, crossing-event and ghosting-event frames, the last two from Stage 5’s relabelling loop).

mAP50 / mAP50-95, precision/recall per class, confusion between danio_rerio and reflection if any, comparison across model versions logged in the registry (e.g. the v3 champion referenced in later stages).

Honest read of where the detector is strong (clean single-fish frames) and where it struggles (the crossing/merge case that motivates Stage 4’s tracker-side handling and evaluation) — the same “fix is in the labels” thread that Stage 4 concludes with. What would improve mAP further if pursued (more crossing frames via the Stage 5 relabelling loop, more morph diversity).