ReferTrack-Qwen3-4B
Checkpoint package associated with ReferTrack: Referring Then Tracking for Embodied Visual Tracking.
This release provides a Qwen3-4B checkpoint for ReferTrack-style embodied visual tracking, intended to align with the official paper, project page, codebase, and demo video.
Official Resources
- Paper: arXiv 2607.20061
- Project page: ReferTrack website
- Code: MedlarTea/referTrack
- Video: YouTube demo
About ReferTrack
ReferTrack is a referring-then-tracking paradigm for embodied visual tracking. The model first grounds a language-described target to an indexed image-space bounding box, then predicts tracking waypoints conditioned on that explicit grounding decision. The method further uses temporal-viewpoint-bbox indicator (TVBI) tokens to preserve target motion cues over time and improves target identification through co-training with Refer-QA data.
EVT-Bench Results
The official paper reports the following single forward-view results on EVT-Bench (Table 1), where each split is evaluated with Success Rate (SR), Tracking Rate (TR), and Collision Rate (CR):
| Split | SR | TR | CR |
|---|---|---|---|
| STT | 89.4 | 92.5 | 1.6 |
| DT | 73.3 | 81.8 | 7.6 |
| AT | 74.1 | 85.7 | 7.7 |
These numbers correspond to the official ReferTrack results reported for the Qwen3-4B single-view setting in the paper.
Files
| File | Description |
|---|---|
refertrack_qwen3_4b.pt |
Model weights (final, stage 2) |
model_config.json |
Model structure and task settings |
refertrack_qwen3_4b_stage1.pt |
Stage-1 weights for training warm-start only (~139 MB; not an inference checkpoint) |
Stage-1 Weights (Training Warm-Start)
ReferTrack is trained in two stages:
- Stage 1 (alignment): the Qwen3-4B LLM, planner and action token are frozen; only the vision projector (
proj) and the TVI embedder (tvi) are trained on color QA for one epoch. - Stage 2 (ReferTrack): full fine-tuning on EVT-Bench refer-navigation and SYNTH-PEDES refer-QA, producing
refertrack_qwen3_4b.pt.
refertrack_qwen3_4b_stage1.pt is the stage-1 result, stripped to the four modules that stage 2 warm-starts from: proj, tvi, planner and act_token ({"model_state": {...}}, 34 tensors). The LLM is not included: stage 2 initializes it from the original Qwen3-4B weights, and adds the 27 refer special tokens itself. Use it only as --pretrained_ckpt when reproducing training; it cannot be evaluated on its own.
Usage
With the ReferTrack code:
# download (add HF_ENDPOINT=https://hf-mirror.com if huggingface.co is unreachable)
bash scripts/eval/download_ckpt.sh # refertrack_qwen3_4b.pt + model_config.json
bash scripts/eval/download_ckpt.sh --stage1 # also refertrack_qwen3_4b_stage1.pt
# Habitat EVT-Bench evaluation
bash scripts/eval/eval_sim_refer_base.sh
# stage-2 training, warm-started from refertrack_qwen3_4b_stage1.pt
bash scripts/train/train_refer.sh
To load the model in Python:
import torch
from referTrack.eval.load_refer_ckpt import load_refer_model
# model_config.json must sit next to the .pt
model, cfg = load_refer_model("data/logs/ReferTrack-Qwen3-4B/refertrack_qwen3_4b.pt", torch.device("cuda"))
Citation
If you use this checkpoint, please cite ReferTrack:
@article{ye2026refertrack,
title={ReferTrack: Referring Then Tracking for Embodied Visual Tracking},
author={Ye, Hanjing and Zeng, Tianle and Zhang, Jiazhao and Wang, Shaoan and Zhang, Zibo and Situ, Weixi and Zhou, Yuchen and Ling, Yonggen and Zhang, Hong},
journal={arXiv preprint arXiv:2607.20061},
year={2026},
url={https://arxiv.org/abs/2607.20061}
}
- Downloads last month
- 13