← Back to work

GMC-Link

Referring multi-object tracking (RMOT) takes a video and a sentence — "moving cars", "left vehicles which are parking" — and tracks every object the sentence describes. Filmed from a moving car, though, a parked vehicle sweeps across the frame while the car ahead barely moves: the motion in the image is the opposite of what happened on the road.

PYTHONPYTORCHOPENCVCOMPUTER VISION
+9.44
Moving-class HOTA, iKUN host (27.70 → 37.14)
GMC-Link architecture: the host tracker (iKUN) scores each tracklet against the expression. The GMC module turns the tracklet’s boxes into a 12-D motion vector and the expression into a sentence embedding, projects both through a shared MLP, and adds their cosine similarity to the host score before the match decision.
PROBLEM

Existing RMOT methods read motion off bounding-box displacement between frames, which mixes the object’s movement with the camera’s — and the sentences this breaks are exactly the ones about movement. Conventional MOT has compensated for camera motion for years (BoT-SORT, UCMCTrack), but how to carry that over to language matching was largely unexplored.

APPROACH

Four stages bolted onto the host. Shi–Tomasi corners on the road half of the frame, tracked with pyramidal Lucas–Kanade and fitted with RANSAC, give a frame-to-frame homography. Accumulated over gaps of 2, 5 and 10 frames, it predicts where a static object would appear; subtracting that ego-motion velocity from the observed one leaves the residual, and the residuals at three gaps plus box geometry form a 12-D motion feature. That feature and a frozen MiniLM sentence embedding are projected into a shared space and trained with InfoNCE. Their cosine similarity, weighted by expression class, is added to the host’s score — its detector, tracker and encoders stay untouched.

RESULT

Pooled HOTA went up on all four host settings (iKUN, TransRMOT, and FlexHook on Refer-KITTI V1 and V2). On moving-class expressions with iKUN as host, HOTA rose from 27.70 to 37.14 (+9.44 ± 0.92 over three seeds); removing either the ego-motion compensation or the multi-scale gaps cuts the gain significantly (p < 0.01). The gain shrinks as the host’s own motion understanding improves — +9.44 on iKUN, +0.43 on FlexHook V1. The module itself runs at 149 FPS on a CPU.

Three frames (85, 94, 100) from a car-mounted camera on a straight road. A parked car on the left, boxed and labelled "Left vehicles which are parking", slides from centre-left to the left edge; the car directly ahead, boxed and labelled "Vehicles in front of us", stays on the centre line in all three.
Refer-KITTI sequence 0005, frames 85, 94 and 100. The parked car (left box) sweeps across the image as the camera passes; the moving car ahead (centre box) barely shifts. Raw image motion gets both wrong; the residual after ego-motion compensation separates them from the same host score.