Skip to content

Stage 2 · Annotation

Manual fish annotation: drawing bounding boxes and storing them as queryable relational records.

Govinda Lienart · AquaMind · Stage 2

Label Studio (Docker) · YOLO export · 4,507 boxes on 837 frames · MySQL 8.4

Abstract

Stage 2 turns the frame dataset from Stage 1 into a supervised-learning dataset in three steps: the frames are uploaded to Label Studio and every visible fish is boxed by hand, the labelled project is exported in YOLO format, and each box is parsed into a relational annotations table keyed back to its source frame. Each labelling batch is also recorded in an annotation_sets table, together with the sampling and Label Studio details it came from, so every box stays traceable to its frame, its batch and its video. The database now holds 4,507 boxes on 837 frames across seven labelling batches, with no orphaned rows and no coordinates outside the normalised range. This is the labelled-data layer the detector is trained on in Stage 3.

A detector is only as good as its labels. For AquaMind, the labels are bounding boxes drawn by hand around every visible fish on the frames extracted in Stage 1, and they are a key input to the detector trained in Stage 3.

Labelling tools export these boxes as a folder of text files, one per image. That format is enough to train a model once, but it forgets where the labels came from: which video, which sampling of frames, which labelling session. Here that matters, because frames come from several sources, labels are added in rounds, and a model has to be retrained on chosen subsets. Without that link it is not possible to say which data a given model saw, or to rebuild a dataset later.

Stage 2 therefore treats labels as data with a history. Each box is stored as a record linked to its frame, and each labelling round is recorded as a batch, so datasets are assembled by query instead of by copying folders. This chapter describes how the frames were labelled and stored, and Stage 3 then trains the detector on the result.

Produce a hand-labelled dataset, stored in a relational database, in which every box is traceable to its frame, labelling batch and video, as the source for training the object detector in Stage 3.

Frames go into Label Studio and come out as YOLO label files, one per frame, each line describing one bounding box. store_annotations.py combines those files with sidecar files that record how the batch was built, and writes two MySQL tables: one that summarises each round of labelling for traceability, and one that holds every individual bounding box. Fig. 1 shows the full pipeline below.

Stage 2 annotation and label-storage pipeline. The frame-indexed dataset from Stage 1 (JPG frames on disk plus rows in the MySQL frames table) is uploaded by upload_labelstudio.py into a new Label Studio project, running in Docker, where every visible fish is boxed by hand. download_labelstudio.py exports the labelled project in YOLO format, producing one text file per frame with one line per bounding box (class, x centre, y centre, width, height, normalised between 0 and 1). store_annotations.py parses each label file, looks up the matching frame_id in the frames table (the Stage 1 table, an input to the script), reads the sidecar files (extraction_params.yaml and download_params.yaml) for provenance, and writes two kinds of rows to MySQL: one row per labelling batch in annotation_sets, and one row per bounding box in annotations, foreign-keyed to both its frame and its batch. Both outputs converge into a labelled dataset in which every box is traceable to its frame and its labelling batch, feeding forward into Stage 3's YOLO training.
FigStage 2 pipeline: frames from Stage 1 are uploaded by upload_labelstudio.py and labelled by hand in Label Studio (the only human step), exported in YOLO format, and stored by store_annotations.py as one annotation_sets row per labelling batch plus one annotations row per bounding box, each foreign-keyed back to its frame. The sidecar files written by the earlier scripts record where each batch came from. Cylinders mark MySQL tables, boxes mark scripts, tools and files. Click to zoom.

The subsections below follow the pipeline figure from top to bottom: Label Studio is set up first, then each script in turn, with the tables introduced right before the script that first writes to them.

Label Studio runs as a Docker container, so nothing is installed locally beyond Docker Desktop. A volume mount keeps its data in a local mydata folder on the host, so labelled work survives a container restart. The upload and download scripts talk to it through its API, using the address and key stored in the .env file.

Label Studio labelling view of a tank frame. Bounding boxes are drawn around each fish, the task list is on the left, the list of regions with their classes is on the right, and the two class buttons, danio_rerio and reflection, sit under the image.
FigLabel Studio, the open-source annotation tool used in this stage. Each task is one frame: boxes are drawn directly on the image, the regions panel on the right lists every box with its class, and the two class buttons under the image (with their hotkeys) are the label set defined by upload_labelstudio.py.

upload_labelstudio.py (flow diagram) reads a project name and a frames folder from config.yaml and creates the project through Label Studio’s API. The labelling interface is defined in the script and not in the web interface: an image with rectangle labels for two classes, danio_rerio (hotkey d) and reflection (hotkey r). Reflections of the fish on the tank glass are a class of their own so that the detector learns to ignore them, instead of never seeing them. The script then uploads every JPG in the folder as one task and logs its progress after each upload.

Each task is opened in Label Studio and a bounding box is drawn around every visible fish, one box per fish, plus one for each visible reflection. This is the only manual step in the stage.

The reflection class is the hard part of the labelling. The glass of the tank acts as a mirror, so a fish can appear twice in the same frame: once as the animal itself and once as its reflection. If both were boxed as fish, one individual would be counted twice, and the detector would be trained to treat a mirror image as a real animal. Boxing reflections under their own class teaches the model the difference. In the figure below, cyan boxes mark real fish (danio_rerio) and pink boxes mark reflections (reflection).

The tank adds a second layer of ambiguity. The glass is not the only reflective surface: the water surface also mirrors the scene, so a reflection can itself be reflected again, and a single fish can leave several look-alike images in the same frame. Some of these are faint, distorted or partly cut off, and it is not always obvious where a real fish ends and a mirror image begins. The labelling therefore needs a consistent rule for what counts as a fish and what counts as a reflection, applied the same way on every frame. Otherwise the two classes blur together in the training data. The aim is that a detector trained on these labels learns to tell them apart, but whether it manages this reliably is something the evaluation in Stage 3 has to show, not something the labelling can guarantee.

Close-up of the left side of the tank in Label Studio. Cyan boxes surround real fish and pink boxes surround their reflections on the glass, several of them overlapping or sitting directly next to the fish they mirror.
FigThe two classes in practice. Cyan boxes are real fish (danio_rerio), pink boxes are their reflections on the tank glass (reflection). Labelling reflections explicitly prevents the same individual from being counted twice and stops the detector from learning to see mirror images as fish.

download_labelstudio.py (flow diagram) looks up the project by name, asks Label Studio for the ids of the tasks that have been labelled, exports them with Label Studio’s YOLO export type, and unpacks the archive into a timestamped folder. Source images can be downloaded as well, but by default they are not, because the frames already exist on disk from Stage 1. Next to the export the script writes download_params.yaml, which records the project name and id, the task-id range and the download time.

The export contains a classes.txt file with the label names, a labels/ folder with one .txt file per frame, and a notes.json file with metadata. Each label file has one line per bounding box, holding five values that are normalised between 0 and 1 relative to the image dimensions, which makes the format independent of resolution. For example:

class_id x_center y_center width height
0 0.242 0.668 0.103 0.113

This line is one fish: class 0 (danio_rerio), centred about a quarter of the way across the frame and two thirds of the way down.

Diagram of one fish in a tank image with its YOLO bounding box. The label line 0 0.242 0.668 0.103 0.113 is broken down into the class (danio_rerio), the x and y centre of the box, and its width and height, each normalised between 0 and 1.
FigYOLO coordinate system: the label line for one fish, broken down. All values are normalised between 0 and 1.

Stage 2 adds two relational tables, both first written to by the script below. annotations holds one row per bounding box: the frame it sits on (frame_id, a foreign key into frames), the labelling batch it came from (annotation_set_id, a foreign key into annotation_sets), the class, and the four normalised box coordinates. annotation_sets holds one row per labelling batch: the video (video_id), the frame_source (regular, crossing_event or ghosting_event), the sampling parameters (sample_rate, start_seconds, end_seconds, frames_extracted, plus iou_threshold and dedup_window for crossing-event batches), and the Label Studio project details (name, id, download time and task-id range).

The batch table was added later in the project, once frames came from more than one source (see Stage 5), and annotations gained its annotation_set_id column at the same time. Keeping the batch details in their own table means they are recorded once per batch, not repeated on every box, the same split Stage 1 uses for videos and frames.

Entity-relationship diagram of all four MySQL tables. Stage 1 tables: videos (one row per registered video, with tank and recording metadata) and frames (one row per extracted frame; video_id is a foreign key into videos.id, with a unique constraint on video_id and frame_number). Stage 2 tables: annotation_sets (one row per labelling batch; video_id is a foreign key into videos.id, with frame_source, sampling parameters and the Label Studio project details) and annotations (one row per bounding box; frame_id is a foreign key into frames.id and annotation_set_id is a foreign key into annotation_sets.id, with class_id, label and the four normalised YOLO coordinates). Crow's-foot lines mark four one-to-many relationships: one video has many frames, one video has many annotation sets, one frame has many annotations, and one annotation set has many annotations.
FigEntity-relationship diagram of the four tables of the aquamind database. The two navy tables come from Stage 1, the two teal tables are added in Stage 2. Every bounding box in annotations points to both the frame it sits on (frame_id) and the labelling batch it came from (annotation_set_id), and both of those trace back to a video. Crow's-foot notation: bar for one, three prongs for many. Click to zoom.

The definitions of all four tables are in schema.sql on GitHub.

store_annotations.py (flow diagram) is where the three kinds of input in the pipeline figure meet: the exported label files, the frames table from Stage 1, and the two sidecar files. It first reads the sidecars: extraction_params.yaml, written by Stage 1 when the frames were extracted, and download_params.yaml, written earlier in this stage by download_labelstudio.py. From the two combined, plus the video name in config.yaml (used to look up the video’s id), it creates one annotation_sets row for the batch, which gives it a batch id. It then loops over the .txt files in labels/, one per frame.

Label Studio prefixes each exported file with a hash (for example e6d83681-frame_360.txt), so the script parses the frame number out of the file name and looks it up in the frames table to get the frame_id. The lookup matches on both the frames folder and the frame number, because the same frame number can occur in different extraction runs of the same video. A frame that is not in the table is skipped with a warning, so a label can never be stored against a missing frame.

Each line in a label file is one box. The five values are cast (class_id to int, the four coordinates to float), a line with fewer than five values is skipped with a warning, and the class id is mapped to its name. Every box is inserted into annotations with its frame_id and the batch id, and a single commit at the end saves the whole batch, so a run that fails partway saves nothing. The script finishes by printing the batch id and a ready-made SQL query for counting the boxes in it.

Correctness was verified by querying the tables directly (counts as of 2026-09-21). Seven labelling batches are stored, from three kinds of frame source:

Annotation set Frame source Frames with boxes Bounding boxes
5 regular 46 445
8 crossing_event 121 781
9 regular 103 719
10 regular 111 224
11 crossing_event 73 355
14 crossing_event 109 598
17 ghosting_event 274 1385
Total 837 4,507

Frames that were uploaded but ended up with no stored box do not appear in annotations, so the frame counts here are frames that hold at least one box.

Beyond the counts, the stored boxes were checked for a missing batch link and for coordinates outside the normalised range. A missing frame link is ruled out by the foreign key on frame_id. The class balance and the boxes per frame were read off the same tables:

Check Result
Boxes without a labelling batch 0
Boxes outside the normalised range 0
Boxes per class 3,600 danio_rerio, 907 reflection
Boxes per frame 5.4 on average, 26 at most

Stage 2 turns hand-drawn boxes into a queryable dataset. Each box is a row linked to its frame and to the labelling batch it came from, and each batch is linked to its video, so a question such as which frames of which video a model was trained on has a SQL answer. Keeping the batches in their own table is what lets Stage 3 pick a training set by listing batch ids, and it is what makes each model version traceable to the data it saw.

Two design choices are worth stating. Reflections are labelled as their own class instead of being left unlabelled, so the detector is shown what to ignore. And the frame lookup checks the frames folder as well as the frame number, so labels from different extraction runs cannot be attached to the wrong frame.

The main limit is that the labelling is manual and was done by a single annotator, so there is no measure of labelling agreement. The frames also come from one tank and a small group of fish. Later rounds reduce the manual cost in two ways, described in Stage 5: the trained detector pre-labels new frames so they only need correcting, and the frames worth labelling are chosen from the tracker’s failures. The labelled dataset feeds Stage 3, where the detector is trained.