license: apache-2.0
We extract the vision encoder from Pi05 base model to caculate the image tokens for data preprocessing.