Skip to content

Stage 8 · Feeding-strike detection

Feeding-strike detection: classical and deep learning both fail, a simple burst signal still helps

Govinda Lienart · AquaMind · Stage 8 · September 2026

Video: IMG_2349.MOV (4-5 fish, food-injection segment) · 118 labelled strikes

Abstract

A feeding strike is a fast, subtle appearance-and-posture event with no visible prey to anchor on, which makes it a harder computer-vision target than the spatially “loud” behaviours of earlier stages. This stage tests whether a strike can be detected automatically from video, and how far a geometry-only approach gets before appearance information is needed. A geometry-only baseline (logistic regression, random forest, XGBoost on the same four kinematic features) lands at or below chance across all three classifiers (0.56 / 0.54 / 0.49 accuracy), which is the evidence that appearance features are actually required. A CNN-LSTM on image crops (DINOv2 vits14 features into an LSTM, with early stopping) reaches 64 % accuracy, 0.66 precision, 0.68 recall on a balanced test split, but a negative control run over a realistic, continuous stream shows that result does not hold: the model calls a strike on 46 % of ordinary-swimming windows versus 64 % of real-feeding windows, close to indistinguishable from chance. Both approaches land as a benchmarked negative result: the strike-defining cues are not resolvable from centroid tracks or re-ID-oriented crops at this frame rate, a data limitation rather than a fixable modelling gap. A third, deliberately simpler approach then recovers a usable signal from the same tracked data: a group-level burst-rate index, built from nothing more than smoothed speed peaks against a control-only, data-derived threshold, rises from a control mean of 2.94 to a feeding mean of 4.17 events/sec, a roughly 1.4x increase, and tracks real feeding closely enough on direct visual inspection to work as a lab-condition readout where both purpose-built classifiers could not.

A feeding strike is the fast lunge and capture a fish makes when it takes prey. Behavioural ecology treats it as one of the single most informative behavioural cues, especially in larval fish, because a trained observer can read from it the factors acting on the animal. It is also very subtle and hard to score by eye, and automated detection is not established for fish at any real scale in ecology. A detector that worked would be valuable, and in the long run could support routine scoring in the lab.

For computer vision, a feeding strike is a harder target than the behaviours in the previous stage. Chasing, depth preference and other spatially “loud” behaviours have a clear geometric signature: they are defined by where a fish is and how it moves relative to the tank or to another fish. A feeding strike is not. Its signal lives in appearance and posture (a brief body vibration or S-bend and a mouth gape) rather than in position. The prey is a tiny particle, invisible at screen resolution, so there is no visible target to anchor on. The whole event is short and easy to miss when the fish is not in sharp focus.

VidTen seconds from the middle of the food-injection phase, cropped to fish 2 alone (tracker overlay on). Fish 2 covers a wide area of the tank in this window; a strike, when it happens, is a brief interruption to that ordinary swimming, not a separate kind of movement visible from a distance.

This stage tests whether a feeding strike can be detected automatically from video, and how far a geometry-only approach gets before appearance information is needed. The one component that might still carry a geometric signal is the burst, since a strike burst may be faster or sharper than one in normal swimming. Even so, I expect it will not be enough on its own, and that reliable detection needs a model working from the image directly: a CNN extracting appearance and posture per frame, a sequence model learning how they change across the strike. The stage is built in that order: a geometry-only baseline first, then a CNN-LSTM pipeline on image crops.

Stage 8 trains a per-fish feeding-strike classifier from a set of hand-labelled strike events. Two approaches are compared on an identical set of windows: a geometry baseline built from movement alone, and a CNN-LSTM built from the fish’s appearance. The appearance model is the one of interest, because a feeding strike has little geometric signature, and the baseline is there to test that rather than assume it. Fig 1 shows the pipeline end to end; the subsections below follow it in order, and each step is a single script whose output file is the next step’s input.

Stage 8 feeding-strike pipeline. Two upstream inputs (tracks.parquet with per-fish crops from the Stage 4 tracker run, and feeding_labels.xlsx with 118 strikes labelled in Phase A) feed build_feeding_windows_fixed_window.py, which builds fixed 45-frame windows with count-matched negatives from the sand-injection control segment. The windows fork into two pipelines. Left: a geometry-only baseline, train_feeding_classifier.py, running logistic regression, random forest and XGBoost on kinematic features; it lands near chance, which is the evidence that appearance features are needed. Right: a CNN-LSTM fixed-window pipeline (build_feeding_crops_fixed_window.py, then build_feeding_embeddings_fixed_window.py running a frozen DINOv2 vits14 backbone to a per-frame embedding, then train_feeding_lstm_fixed_window.py training an LSTM over the embedding sequence with early stopping), reaching 64 percent accuracy on the balanced test split. That trained model then feeds infer_feeding.py, a negative control that runs it as a sliding-window classifier over full sand and food phases of the same recording: it calls a strike on 46 percent of no-food windows versus 64 percent of real-feeding windows, mean P(strike) 0.48 versus 0.60, showing the balanced accuracy figure does not hold on a realistic stream. Because neither the geometry baseline nor the CNN-LSTM survives that stream test, an arrow leads from the negative-control result down to a fallback box: render_feeding_burst_overlay.py, a purely descriptive per-fish burst-rate metric compared pooled across the control and feeding phases, with no per-strike claim. Accuracy figures are reported in the Results section.
FigStage 8 pipeline: from the reused tracker outputs and the strike labels, through fixed 45-frame windows, to two approaches compared on identical data, a geometry baseline (pink) and the CNN-LSTM (green scripts, orange model). The baseline landing near chance is what motivates the appearance model; the trained model is then stream-tested by infer_feeding.py, a negative control that lands near chance as well. Because neither approach survives that test, it falls back to a descriptive burst-rate metric (purple) rather than another classifier. Performance figures are in Results. Click to zoom.

The feeding-experiment recording (IMG_2349.MOV; watch the tracked recording ↗) is already tracked from earlier stages, so tracks.parquet and the per-fish crops are reused directly. It has three phases of about two minutes each: a calm baseline, a control-injection phase in which tiny sand particles are delivered through the same tubing later used for the food (frames 7505 to 14340), and a food-injection phase (from frame 14340). 118 real strikes were labelled by hand, all of them in the food-injection phase (frames 14617 to 21656).

Positives are taken from the food phase and negatives from the control-injection phase, not from the calm baseline. Both phases involve an injection through the tubing and the fish reacting to it, so the main systematic difference between a positive and a negative window is whether food was present; calm-baseline negatives would instead let the model succeed by detecting disturbance rather than feeding. The match is approximate, since injected sand settles within seconds while a feeding bout keeps the fish foraging for longer, but it is a closer control than calm footage. The chaotic feeding-frenzy burst right at the food-injection point is excluded from both classes: occlusion and ID-switch risk there make reliable per-fish attribution unsafe.

build_feeding_windows_fixed_window.py turns each labelled strike into a fixed-length window centred on the midpoint of its labelled frame range, the frame where the strike itself is visually clearest. A feeding strike, as labelled by eye, is an extremely subtle event: a brief shake or vibration of the body, sometimes with the mouth visibly opening, and it is frequently accompanied by a change of swimming direction right at the strike, as the fish turns onto the food. Any one of these three cues (the shake, the mouth, the turn) can be the clearest signal in a given case, and none of them is reliably visible in every strike, which is part of why the window is centred on the labeller’s best estimate of that moment rather than on a single detectable trigger. A fixed length is used rather than each event’s real span because a deployed detector sees a continuous stream with no event boundaries: a fixed window slides along the input directly, while an event-length window would first need a separate stage to locate candidate events. The length is set to 45 frames (0.75 s at 60 fps), the 75th percentile of labelled strike duration, so most windows contain the whole strike and the rest still contain its central lunge; a longer window was avoided because strikes occur roughly once a second during the feeding bout, so it would often capture a second strike. For each positive window the script draws one non-overlapping negative window at random from the control-injection phase, assigned to a random fish. Any window holding a frame the tracker flagged as occluded or missing is dropped, so both approaches train and test on exactly the same windows. The script then makes a label-stratified 80/20 train/test split, before any crops or features exist; because there is one non-overlapping window per event, no frame from a test strike reaches training. Its output, train_df and test_df (one row per window: fish id, frame range, label, split), is the single source of truth for window membership everywhere downstream.

Speed and burst are used here because they are the only strike cues that survive centroid tracking at all. Of the three visual cues described above, the body shake is a genuine movement event and so has some chance of showing up in a speed trace; the mouth gape does not; the fish is filmed side-on in a whole-tank shot, and a strike can happen facing away from or across the camera as easily as toward it, so the mouth is frequently too small on screen or hidden by the fish’s own body and orientation to be usable as a tracked geometric feature at all, even before considering resolution. The direction change is closer to a heading feature than a speed one, and is not included in this baseline; it is the natural next feature to add if this approach is revisited (see Discussion).

train_feeding_classifier.py reads the window tables and computes four kinematic features per window: mean and maximum speed, and mean and maximum burst, where burst is the frame-to-frame change in speed. Speed is capped at the 99.9th percentile and smoothed with a 5-frame centred rolling mean before burst is derived, to suppress single-frame tracker-jitter spikes. The four features feed three standard classifiers, logistic regression, a random forest and XGBoost, left at their library defaults with no hyperparameter tuning and compared directly: if a real boundary existed in these features, a plain off-the-shelf model should be enough to find it, so a null result here is not explained away by “the wrong algorithm.” Mean and maximum speed, mean and maximum burst are the most direct kinematic quantities a fish’s tracked position can give: the first pair asks how fast it moved, the second how sharply that speed changed, and between them they are the natural first features to try before reaching for anything more elaborate.

Approach 2 — Appearance model (CNN-LSTM)

Section titled “Approach 2 — Appearance model (CNN-LSTM)”

The appearance path is three scripts. build_feeding_crops_fixed_window.py expands each window into 45 ordered per-frame rows, each pointing at that frame’s crop image (the tracker already saved the crops; this only builds the manifest), written grouped by event. build_feeding_embeddings_fixed_window.py loads a frozen DINOv2 vits14 backbone once and passes each window’s 45 crops through it in frame order, producing a (45, 384) embedding sequence; because the backbone never trains, these are computed once and cached as .pt files rather than recomputed each epoch. train_feeding_lstm_fixed_window.py then trains the only part that learns: a single-layer LSTM (hidden size 64) that reads the embedding sequence, with a linear head mapping its final hidden state to strike or no strike. Training uses cross-entropy with Adam (learning rate 1e-3, batch size 8) for up to 100 epochs.

Forward pass of the feeding-strike LSTM. 45 per-frame crops (45x3x224x224) pass through a frozen DINOv2 vits14 backbone into a (45, 384) embedding sequence. An unrolled LSTM reads the sequence one frame at a time; its hidden state carries forward at every step, but each intermediate hidden state is discarded, only the final one (64-dim) is kept. A linear head (64 to 2) maps that vector to two logits, and softmax turns them into a strike / no-strike call.
FigThe model's forward pass on one window: 45 per-frame crops through a frozen DINOv2 backbone, an unrolled LSTM whose intermediate hidden states are computed then thrown away, and a linear head reading only the final one. What happens before and after this exact computation, at training time and at inference time, is described below.

At training time, this forward pass runs once per window inside each batch of 8, followed by cross-entropy loss and a backward pass (Adam, lr=1e-3), repeated for up to 100 epochs, keeping the checkpoint at minimum test loss. At inference time, the identical forward pass runs once per window, with no labels and nothing after it but recording the score; sliding across a continuous video is a separate, outer loop wrapped around this same computation, not a different model.

Both models are first scored on the same held-out 20 percent test split, which is untouched during feature smoothing, training and checkpoint selection. Accuracy, precision and recall are reported for the strike class, where a false positive is a window predicted as a strike that was not one. For the CNN-LSTM, training and test loss are logged every epoch and the curves are read directly rather than trusting the final-epoch score, since on a dataset this small the training loss can keep falling well after the test loss has bottomed out; the model kept for reporting is the checkpoint at minimum test loss, not the last epoch.

A separate script, infer_feeding.py, evaluates the trained CNN-LSTM a second way. Rather than scoring a fixed, curated set of windows, it runs the same forward pass shown in Fig 2 as a sliding-window classifier directly over two full segments of the same recording, the sand-injection phase and the food-injection phase, stepping a 45-frame window forward 20 frames at a time so consecutive windows overlap. Any window missing a tracked crop anywhere in its 45 frames is dropped; every window that survives is scored independently, with no labels, no train/test split and no backpropagation involved, since the model is already trained and frozen at this point. This is a direct test of whether the balanced-split numbers above hold up once the model is run the way it would actually be used: continuously, on ordinary swimming as well as real feeding, rather than on a hand-picked 50/50 set.

Approach 3 — Burst-rate (group-level fallback)

Section titled “Approach 3 — Burst-rate (group-level fallback)”

Once Approaches 1 and 2 both failed the stream test above, render_feeding_burst_overlay.py drops per-fish, per-strike attribution entirely. It reuses the same tracks.parquet and per-fish speed (smoothed the same way as the geometry baseline) to detect bursts as local peaks in speed via scipy.signal.find_peaks, then pools burst counts across all fish rather than assigning any burst to a specific strike or fish. Bursts are compared as a rate (events per second) across two segments of the same recording: control (calm baseline plus the sand-injection phase, collapsed together since both are “no real food”) and real feeding. This is a deliberately smaller claim than either classifier above: it asks only whether the group’s overall burst rate shifts with real feeding, not which fish struck or when.

A speed peak only counts as a burst if it clears three gates, all applied in one find_peaks call:

  • height ≥ BURST_MIN_CMS, the 85th percentile of smoothed speed measured during the control segment alone, with the feeding segment excluded from that calculation so the effect under study cannot leak into the baseline used to detect it (data-derived rather than picked by eye, landing at roughly 8.75 cm/s here)
  • prominence ≥ 4.0 cm/s, so the peak has to stand out from its own local dip, not just ride along the top of an already-elevated stretch
  • distance ≥ 0.4s between peaks, so two spikes closer together than that count as one burst rather than the same lunge counted twice
Density-normalised histogram of per-frame smoothed speed, control segment (baseline plus sand injection, blue) versus feeding segment (orange), with a dashed red line marking the 85th-percentile burst threshold derived from the control segment alone.
FigControl versus feeding speed distributions (density-normalised so the roughly 2x difference in segment length does not distort the comparison). The two curves overlap heavily through the shared peak, so no single frame is classifiable by speed alone, but feeding sits visibly higher through the tail past the threshold, the population-level signal the burst-rate metric below is built on.

Three approaches were tested, in the order the stage was built: a geometry-only baseline, the CNN-LSTM appearance model, and, once neither survived a realistic stream test, a purely descriptive burst-rate fallback. Each is reported here as its own approach, with its own numbers.

All three classifiers are trained on the same four kinematic features (mean speed, maximum speed, mean burst, maximum burst), scored on the strike class, on the same held-out 20 percent test split used for every model in this stage (balanced roughly 50/50 between labelled strikes and matched negatives). Reported separately rather than as one collapsed number, since the point is whether any classifier can find a usable boundary in these features, not which one happens to win:

Model Accuracy Precision (feeding) Recall (feeding)
Logistic regression 0.56 0.55 0.80
Random forest 0.54 0.54 0.70
XGBoost 0.49 0.50 0.45

All three sit at or below chance-equivalent for a balanced binary split, with XGBoost landing exactly at coin-flip accuracy. Logistic regression’s higher recall alongside near-chance precision is a further sign of a weak, mostly linear nudge rather than a real decision boundary: the model is closer to calling “feeding” more often than to actually separating the classes. This is the expected result, not a failure: it is the evidence that a feeding strike has too little geometric signature for movement alone, and that appearance information is needed.

Model Accuracy Precision Recall
CNN-LSTM, fixed 45-frame window ~0.64 0.66 0.68

On the balanced test split the CNN-LSTM looks like it might be a working detector. It is not one.

A negative control makes this concrete. infer_feeding.py runs the trained model as a sliding-window classifier over two full segments of the same recording: the sand-injection phase, where no food is present and the correct strike count is zero for every fish, and the food-injection phase, where real feeding happens.

Segment Positive windows Rate Mean P(strike)
No food (sand-injection phase) 497 / 1080 46.0% 0.476
Real feeding 587 / 917 64.0% 0.598

The model calls a strike on nearly half the windows drawn from ordinary swimming with no food present, against a true count of zero, and on 64 percent of windows during real feeding, only 18 points higher. The 64 percent balanced-split accuracy figure is real; it was measured on the wrong population.

The LSTM’s own decision-threshold sweep, recorded before the negative control was ever run, already carried a warning. Raising the classification threshold from 0.5 to 0.6 to 0.7 made precision, recall and accuracy all worse at once, with 0.5 winning outright. A model with real class separation trades recall for precision as the threshold rises, keeping only its most confident calls; this one had no such structure to isolate, because its output probabilities were never meaningfully separated in the first place, the same fact the stream test later confirmed directly (0.476 vs 0.598, both close to the 0.5 midpoint).

Approach 3 — Burst-rate (group-level fallback)

Section titled “Approach 3 — Burst-rate (group-level fallback)”

Neither the geometry baseline nor the CNN-LSTM survives a realistic stream, so the last approach drops per-fish, per-strike attribution altogether and asks a smaller, honest question instead: does a fish’s burst rate change between control and feeding at the group level? render_feeding_burst_overlay.py counts speed peaks (bursts) per fish from the same tracked recording, pooled across all fish, with no attempt to say which fish struck or when.

Segment Pooled burst rate (events/sec)
Control (baseline + sand injection) 2.94
Real feeding 4.17
Line plot of pooled burst rate in 10-second bins over the full recording, shaded blue for the control segment and orange for feeding, with dashed horizontal lines at each segment's mean (control 2.94/s, feeding 4.17/s).
FigBurst rate over time, control versus feeding, binned every 10 seconds and pooled across all fish. The rate visibly steps up at the control/feeding boundary and stays elevated for the rest of the recording, rather than spiking once and decaying back to baseline.
VidThe full recording with the burst overlay live: a star flashes on a fish the instant its speed crosses the burst threshold, a running per-fish control/feeding count updates in the corner, and a BASELINE / SAND (control) / FEEDING subtitle at tank-floor level marks which experimental phase is currently playing. The counters visibly accelerate once the video crosses into the feeding segment.

Burst rate rises roughly 1.4x from control to real feeding. It does not need to solve the per-strike attribution problem the previous two tests failed at, since it never claims to say which fish struck or when; it is a legitimate before/after readout for the alarm-cue assay, and the closest thing to a working deliverable this stage produced.

A model that flags a strike on nearly half of all ordinary swimming is not a slightly-off detector. It is not a detector at all, and that is the real headline of this stage, not the 64 percent balanced-accuracy score that looked like a working result until it met a realistic stream instead of a curated 50/50 split. The same trained model called a strike on 46 percent of ordinary-swimming windows, true count zero, barely below its 64 percent hit rate on windows that actually held one. A number that confident and that wrong is not a tuning problem, it traces to what the data can carry, not to the model or the pipeline. A zebrafish strike is defined by a mouth gape and eye convergence, at a scale specialist studies (Shamur et al. 2016, filming at 240 fps with head-orientation normalisation) still cap around 73 percent accuracy for. At 60 fps, whole-tank, working from centroid tracks and re-ID-oriented crops, neither cue is resolvable, confirmed directly on a labelled clip where the fish’s centroid barely moves during a strike, since the event is a head snap rather than a translation.

Studies that do recover gape and posture at this level of detail, such as Mearns et al. (2020), classify capture bouts from tail-posture and timing-to-max-gape measurements extracted at a level of kinematic detail this whole-tank, 60 fps setup was never built to capture. The plausible mechanism is that a strike this brief and this small on screen leaves little trace once compressed into a bounding-box crop and an embedding, though this pipeline did not isolate that mechanism directly, only its consequence. More labelled data would not fix it either way: both approaches failed from a ceiling on what the features themselves carry, not from too few examples, so what would actually help is different data, a higher frame rate and a jaw or head keypoint, not more labels from the same recording. Phase G (an overlay, per-fish counters, an automated report) is dropped accordingly; there is nothing trustworthy to visualise from either classifier, and the chasing classifier from the previous stage remains this project’s working behaviour detector.

The clearest surprise of this stage sits outside the classifiers entirely. Watching the burst-rate index (Approach 3) directly against the video by eye, it tracks real feeding closely: minutes with visibly more fast movement are, on inspection, minutes where fish are actually striking, closer than either learned model ever managed on its own test split. The likely reason is the same failure seen from the other side. The CNN-LSTM was asked to learn a cue that is not present in its inputs at this resolution, so its output probabilities sit near 0.5 and flicker close to at random, while peak detection on smoothed speed asks a much smaller, answerable question, not “is this fish gaping at prey” but “is this fish moving unusually fast right now.” Complexity did not help here because the extra capacity had nothing extra to learn from, and the simplest tool in the stage ended up reading the behaviour best.

The match is not unconditional: as food is consumed and becomes scarce, bursts blend back into ordinary locomotor activity and can no longer be told apart from a fast turn on their own. Read that decline as a feature rather than a flaw and the metric is doing something biologically sensible, tracking perceived prey availability rather than strikes directly, which is exactly the behaviour a foraging fish should show as a patch depletes. What survives for measurement is the trend around a known injection point, not the individual event, a strong lab-condition tool rather than a general in-vivo one. That trend is still only as trustworthy as the cutoff used to call a speed peak a burst in the first place, and the control and feeding distributions overlap enough that no single frame gives away which condition it came from; somewhere past their shared peak, in the population of frames rare rather than typical, the two curves finally pull apart.

The threshold used here is the 85th percentile of speed from the calm-baseline-plus-sand segment, with the food-injection segment excluded from that calculation on purpose. The exclusion matters more than the percentile: fold feeding speeds into the same calculation and a real speed increase from feeding raises the very bar feeding then has to clear, so a strong effect would make itself harder to see, not easier. Someone has already tried removing that kind of arbitrariness altogether. Wang, Kroeschell, Zhu, Mumm and Yi (bioRxiv, 2026) train a CNN-BiLSTM on hand labelled tail kinematics to detect swim bouts in tethered larval zebrafish with no cutoff at all. Their own numbers mark where that effort earns its keep: on high-amplitude bouts the learned model ties a plain threshold, 97.7 against 98.0 percent accuracy, and only pulls ahead on subtle, low-amplitude events, where precision climbs from 67 to 89 percent. A feeding-driven burst is the large, high-amplitude case their results already call solved without a model, which is why a percentile threshold, not a learned one, was kept here.

What that threshold cannot do yet is travel. Run this same calibration on a different day, tank, or the same fish a month later, and each one quietly recomputes its own cutoff, blind to every other session, so a genuine rise in baseline activity across a study would vanish into a moving target rather than showing up against a fixed one. The fix costs nothing extra: pool control-segment speed across every session or tank first, take the percentile once, and hold it fixed for the rest of the study, a number that means the same thing on day one and day thirty, tank A and tank B, before anyone outside a single recording has reason to trust it. That is the natural next step for per-strike detection too, treated as still open rather than closed: a movement-based feature set richer than raw speed, closer to the bearing and lagged path-following features already scoped for the chasing classifier in Stage 7, combined with the burst signal and its pooled threshold, is a more promising direction than a bigger appearance model on the same crops. None of that is built or claimed here; it is the concrete direction this negative result points toward, needing considerably more labelled sessions and cross-tank testing before any of it could be trusted.

Shamur, E., Zilka, M., Hassner, T., China, V., Liberzon, A., & Holzman, R. (2016). Automated detection of feeding strikes by larval fish using continuous high-speed digital video: a novel method to extract quantitative data from fast, sparse kinematic events. Journal of Experimental Biology, 219(11), 1608-1617.

Wang, H., Kroeschell, G., Zhu, Y., Mumm, J. S., & Yi, J. (2026). Threshold-free neural network models for swim bout detection of larval zebrafish. bioRxiv.

Mearns, D. S., Donovan, J. C., Fernandes, A. M., Semmelhack, J. L., & Baier, H. (2020). Deconstructing hunting behavior reveals a tightly coupled stimulus-response loop. Current Biology, 30(1), 54-69.