VisionHOPE ONNX
ONNX exports of PSRben/VisionHOPE (paper, code) for onnxruntime / onnxruntime-web (WebGPU). They power the in-browser demo at mrfakename/visionhope-webgpu.
Files (onnx/)
| file | task | input | outputs |
|---|---|---|---|
visionhope_{tiny,small,base}_cls.onnx |
ImageNet-1k classification | pixel_values float32 [1,3,224,224] |
logits [1,1000] |
visionhope_{tiny,small,base}_ade20k_512.onnx |
ADE20K semantic segmentation (UPerNet) | pixel_values float32 [1,3,512,512] |
logits [1,150,128,128], labels uint8 [1,512,512] |
*_parity.json next to each model holds the parity numbers below.
Static shapes, opset 17, fp32. Each graph was exported from a pure-PyTorch port of the official model (the SRNL scan is unrolled over row/column chunks), then pre-optimised with onnxruntime's EP-independent basic level (constant folding), so sessions start in seconds. Detection (Mask R-CNN) is not exported: its NMS and dynamic RoI stages do not map well to a static WebGPU graph.
Preprocessing
- Classification: RGB, resize short side to 224 (bicubic in the reference), center-crop 224, scale to [0,1], normalize with mean (0.485, 0.456, 0.406), std (0.229, 0.224, 0.225), NCHW.
- Segmentation: RGB resized to 512×512 (aspect ratio not kept), normalize in 0–255 space with
mean (123.675, 116.28, 103.53), std (58.395, 57.12, 57.375), NCHW.
labelsis the argmax after bilinear upsampling to 512×512; resize it back to the image size with nearest neighbour. This differs from the official evaluation (keep-ratio, short side 512, sliding/whole inference), so scores will not match the paper exactly.
Parity (ONNX Runtime CPU vs. PyTorch)
| model | nodes | size | images | agreement | max abs diff (logits) | ORT CPU s/img |
|---|---|---|---|---|---|---|
visionhope_tiny_cls.onnx |
52,889 | 129 MB | 6 | top-1 6/6, top-5 6/6 | 1.9e-05 | 0.9 |
visionhope_small_cls.onnx |
80,510 | 248 MB | 6 | top-1 6/6, top-5 6/6 | 2.8e-05 | 1.7 |
visionhope_base_cls.onnx |
80,510 | 407 MB | 6 | top-1 6/6, top-5 6/6 | 3.7e-05 | 2.9 |
visionhope_tiny_ade20k_512.onnx |
128,179 | 278 MB | 4 | pixels ≥ 99.9996% | 4.0e-05 | 4.2 |
visionhope_small_ade20k_512.onnx |
195,050 | 420 MB | 4 | pixels ≥ 99.9996% | 4.8e-05 | 6.1 |
visionhope_base_ade20k_512.onnx |
195,050 | 589 MB | 4 | pixels 100% | 4.8e-05 | 7.4 |
Reference = the PyTorch port in fp32 on the same inputs (COCO val2017 images; 6 for classification, 4 for segmentation), onnxruntime 1.20.1 CPU on 16 vCPUs. Segmentation agreement is the share of identical argmax labels on the 512×512 map; the abs diff is on the 128×128 logits. All six models were also run in headless Chrome with onnxruntime-web on WebGPU (see the Space).
The PyTorch port itself was validated against the official implementation (mmcv/mmseg/timm, CUDA kernels): classification logits within 1e-5, ADE20K pixel agreement ≥ 0.99998.
Usage (Python)
import onnxruntime as ort, numpy as np
s = ort.InferenceSession("visionhope_tiny_cls.onnx")
logits = s.run(None, {"pixel_values": x.astype(np.float32)})[0]
Model tree for mrfakename/VisionHOPE-ONNX
Base model
PSRben/VisionHOPE