← 回文章列表

Building Refer-MOT: When Language Meets Vision

A deep dive into how I combined ByteTrack with a Vision-Language Model to achieve a 53.5% HOTA score improvement on the Refer-KITTI benchmark.

The Problem

Multi-Object Tracking (MOT) has traditionally been a purely visual task — detect objects, associate them across frames, and maintain consistent IDs. But what if you want to track specific objects described in natural language? That’s the challenge of Referring MOT.

Imagine a self-driving car receiving the command: “Follow the red sedan that just turned left.” The system needs to understand both the visual scene and the linguistic reference to identify and track the correct vehicle.

My Approach

I built a multimodal pipeline that bridges two powerful systems:

1. ByteTrack for Robust Tracking

ByteTrack is a simple yet effective MOT algorithm. Instead of discarding low-confidence detections like most trackers, it uses them as a second association pass. This dramatically reduces ID switches and handles occlusion gracefully.

2. Vision-Language Model for Semantic Grounding

I integrated a VLM to ground natural language descriptions onto detected objects. The model takes a text query and a set of detection crops, then scores each crop based on semantic similarity.

The Pipeline

Input Frame → Detector → ByteTrack Association
                              ↓
                    Track Candidates
                              ↓
           VLM Scoring (text query × crop embeddings)
                              ↓
                 Filtered Referred Tracks

Key Challenges

Distribution Mismatch: The VLM was trained on clean, well-cropped images, but ByteTrack feeds it noisy, partially occluded crops. I had to implement a confidence-weighted scoring mechanism to handle this.

Temporal Consistency: A single-frame VLM score can fluctuate wildly. I added a temporal smoothing window that averages scores across 5 frames, dramatically reducing false positives.

Speed: Real-time performance was non-negotiable. I optimized the VLM inference path with batched processing and cached embeddings for the text query.

Results

The final system achieved a +53.5% HOTA score boost on the Refer-KITTI benchmark. HOTA (Higher Order Tracking Accuracy) is a comprehensive metric that balances detection accuracy, association accuracy, and localization quality.

Lessons Learned

  1. Simple baselines first — ByteTrack’s simplicity made it easy to debug and extend
  2. Distribution matters — always validate your model on data that looks like what it’ll actually see
  3. Temporal context is free signal — if you’re processing video, use the time dimension

The code is open-source on GitHub. Feel free to explore, fork, or reach out with questions.