← Back to work

Refer-MOT

Ordinary multi-object tracking is purely visual — detect, associate across frames, keep IDs stable. Refer-MOT has to track the object a sentence describes, which means the tracker has to read language.

PYTHONPYTORCHOPENCVVLM
+53.5%
HOTA boost on Refer-KITTI
Architecture diagram (drop your image here)
PROBLEM

A purely visual tracker cannot tell which target a description refers to, and bolting the language condition on as post-processing leaves semantics and trajectories out of step.

APPROACH

A multimodal pipeline: ByteTrack owns the trajectories, a vision-language model does the semantic reasoning, and a visual-question-answering verification loop crops candidate objects on the fly to confirm each one actually matches the prompt.

RESULT

HOTA on Refer-KITTI rose from 14.08 to 21.62 — a 53.5% relative gain. GMC-Link grew out of the same line of work.