ReferTrack-Qwen3-4B

Checkpoint package associated with ReferTrack: Referring Then Tracking for Embodied Visual Tracking.

This release provides a Qwen3-4B checkpoint for ReferTrack-style embodied visual tracking, intended to align with the official paper, project page, codebase, and demo video.

Official Resources

About ReferTrack

ReferTrack is a referring-then-tracking paradigm for embodied visual tracking. The model first grounds a language-described target to an indexed image-space bounding box, then predicts tracking waypoints conditioned on that explicit grounding decision. The method further uses temporal-viewpoint-bbox indicator (TVBI) tokens to preserve target motion cues over time and improves target identification through co-training with Refer-QA data.

EVT-Bench Results

The official paper reports the following single forward-view results on EVT-Bench (Table 1), where each split is evaluated with Success Rate (SR), Tracking Rate (TR), and Collision Rate (CR):

Split SR TR CR
STT 89.4 92.5 1.6
DT 73.3 81.8 7.6
AT 74.1 85.7 7.7

These numbers correspond to the official ReferTrack results reported for the Qwen3-4B single-view setting in the paper.

Files

File Description
refertrack_qwen3_4b.pt Model weights (final, stage 2)
model_config.json Model structure and task settings
refertrack_qwen3_4b_stage1.pt Stage-1 weights for training warm-start only (~139 MB; not an inference checkpoint)

Stage-1 Weights (Training Warm-Start)

ReferTrack is trained in two stages:

  1. Stage 1 (alignment): the Qwen3-4B LLM, planner and action token are frozen; only the vision projector (proj) and the TVI embedder (tvi) are trained on color QA for one epoch.
  2. Stage 2 (ReferTrack): full fine-tuning on EVT-Bench refer-navigation and SYNTH-PEDES refer-QA, producing refertrack_qwen3_4b.pt.

refertrack_qwen3_4b_stage1.pt is the stage-1 result, stripped to the four modules that stage 2 warm-starts from: proj, tvi, planner and act_token ({"model_state": {...}}, 34 tensors). The LLM is not included: stage 2 initializes it from the original Qwen3-4B weights, and adds the 27 refer special tokens itself. Use it only as --pretrained_ckpt when reproducing training; it cannot be evaluated on its own.

Usage

With the ReferTrack code:

# download (add HF_ENDPOINT=https://hf-mirror.com if huggingface.co is unreachable)
bash scripts/eval/download_ckpt.sh            # refertrack_qwen3_4b.pt + model_config.json
bash scripts/eval/download_ckpt.sh --stage1   # also refertrack_qwen3_4b_stage1.pt

# Habitat EVT-Bench evaluation
bash scripts/eval/eval_sim_refer_base.sh

# stage-2 training, warm-started from refertrack_qwen3_4b_stage1.pt
bash scripts/train/train_refer.sh

To load the model in Python:

import torch
from referTrack.eval.load_refer_ckpt import load_refer_model

# model_config.json must sit next to the .pt
model, cfg = load_refer_model("data/logs/ReferTrack-Qwen3-4B/refertrack_qwen3_4b.pt", torch.device("cuda"))

Citation

If you use this checkpoint, please cite ReferTrack:

@article{ye2026refertrack,
  title={ReferTrack: Referring Then Tracking for Embodied Visual Tracking},
  author={Ye, Hanjing and Zeng, Tianle and Zhang, Jiazhao and Wang, Shaoan and Zhang, Zibo and Situ, Weixi and Zhou, Yuchen and Ling, Yonggen and Zhang, Hong},
  journal={arXiv preprint arXiv:2607.20061},
  year={2026},
  url={https://arxiv.org/abs/2607.20061}
}
Downloads last month
13
Video Preview
loading

Paper for hjyeee/ReferTrack-Qwen3-4B