Skip to content
Levijoy

How it works

From raw footage to a jump list — here's what's actually happening.

Levijoy doesn't watch the video the way a person does. It tracks a few points on the horse's body through every frame, turns that into a simple motion signal, and then works out which moments in that signal are real jumps. Here's the pipeline, step by step.

Real competition footage is messy.

Broadcast and phone footage of a jumping round isn't clean lab data. The camera pans and zooms to follow the horse, other riders wander in and out of frame, and jump wings, flowers, and banners clutter the background.

Any system that just looks for "something moving up and down" gets confused fast. Levijoy has to separate the horse's actual jump from everything else happening in the arena.

A competition frame annotated to show the rider, camera motion, obstacle clutter, and another horse in frame

Track the horse, not just detect it.

The first version of this system tried to outline the horse in every frame, the same way a lot of computer vision tools work. That fell apart the moment the horse jumped — the outline fragmented and camera motion swamped the signal.

Levijoy instead tracks three fixed points on the horse's body — the withers, the base of the tail, and the poll (just behind the ears) — through every frame. Following specific points turned out to be far more reliable than trying to outline the whole animal.

Comparison of a fragmented outline mask versus three stable tracked points on a jumping horse

Trained to recognize horses from any angle.

Competition footage isn't filmed from one consistent angle. The camera might catch a round side-on, from an angle, from behind, or with the horse briefly hidden behind a fence.

The tracking model was trained on real footage covering all of these cases, then refined further by finding the frames it struggled with most and training on those specifically — so it keeps working as the camera and horse move around each other.

Four example frames showing the tracking model working from side, oblique, rear, and partially occluded views

Turning motion into a signal.

Once the three points are tracked, Levijoy follows a fixed sequence to go from raw pixel positions to a clean, comparable signal:

  1. Raw frame

    Start with the ordinary video frame — nothing special about the footage is required.

  2. Pose extraction

    Find the withers, tail base, and poll in that frame.

  3. Composite Y

    Track how high off the ground those points sit, frame by frame, picking whichever point is being tracked most reliably at each moment.

  4. Detrended

    Subtract out slow camera panning, leaving only fast up-and-down motion — jumps included.

  5. Normalized

    Adjust for camera zoom, so a jump filmed close-up and one filmed from far away look the same to the system.

  6. Detected jumps

    Find the peaks in that cleaned-up signal — each one is a jump candidate.

The whole pipeline at a glance.

Six-panel diagram showing the raw frame, pose extraction, composite Y position, detrended signal, normalized signal, and detected jumps
From raw frame to detected jumps, in six panels.

Double-checking before confirming a jump.

Not every spike in the signal is a jump. Regular canter strides create smaller, regular bumps, and occasionally the tracking itself glitches — for instance if the model briefly loses the horse in a crowd.

Levijoy filters out canter strides by comparing how prominent each bump is, then double-checks the remaining candidates by fitting a smooth curve to the motion: a real jump arcs cleanly up and back down, while a tracking glitch looks erratic. Anything that doesn't look like a real jump gets rejected before it ever reaches the final count.

Two graphs comparing the smooth arc of a real jump against the erratic shape of a tracking glitch

Continuously learning from real footage.

The tracking model wasn't trained once and left alone. After the first round of training, it was run on footage it hadn't seen, and the frames it struggled with most were fed back in as additional training data.

That feedback loop meaningfully improved how precisely the model locates each point on the horse — and it's the same process we'll keep using as Levijoy sees more yards, horses, and camera setups.

Two-stage training diagram: initial training on 14 videos, then failure analysis on unseen footage, about 80 relabelled difficult frames, and retraining that cuts the tracking error from 28.27 px to 14.74 px

Want to see this running on real footage?