GMC-Link
Referring multi-object tracking (RMOT) takes a video and a sentence — "moving cars", "left vehicles which are parking" — and tracks every object the sentence describes. Filmed from a moving car, though, a parked vehicle sweeps across the frame while the car ahead barely moves: the motion in the image is the opposite of what happened on the road.

Existing RMOT methods read motion off bounding-box displacement between frames, which mixes the object’s movement with the camera’s — and the sentences this breaks are exactly the ones about movement. Conventional MOT has compensated for camera motion for years (BoT-SORT, UCMCTrack), but how to carry that over to language matching was largely unexplored.
Four stages bolted onto the host. Shi–Tomasi corners on the road half of the frame, tracked with pyramidal Lucas–Kanade and fitted with RANSAC, give a frame-to-frame homography. Accumulated over gaps of 2, 5 and 10 frames, it predicts where a static object would appear; subtracting that ego-motion velocity from the observed one leaves the residual, and the residuals at three gaps plus box geometry form a 12-D motion feature. That feature and a frozen MiniLM sentence embedding are projected into a shared space and trained with InfoNCE. Their cosine similarity, weighted by expression class, is added to the host’s score — its detector, tracker and encoders stay untouched.
Pooled HOTA went up on all four host settings (iKUN, TransRMOT, and FlexHook on Refer-KITTI V1 and V2). On moving-class expressions with iKUN as host, HOTA rose from 27.70 to 37.14 (+9.44 ± 0.92 over three seeds); removing either the ego-motion compensation or the multi-scale gaps cuts the gain significantly (p < 0.01). The gain shrinks as the host’s own motion understanding improves — +9.44 on iKUN, +0.43 on FlexHook V1. The module itself runs at 149 FPS on a CPU.
