1qh/chandra-ocr-2-oq4-mlx

datalab-to/chandra-ocr-2 quantized with oQ (oMLX v0.5.3) at level 4.0, data-driven mixed precision: oQ measures each layer's quantization error through calibration and allocates bits where the measurement says they matter, rather than by a fixed per-tensor rule.

Base datalab-to/chandra-ocr-2
oQ level 4.0 (effective ~4.6–4.7 bits/weight)
Size 3.2 GB (from 9.1 GB bf16)
Sensitivity entries 248
oMLX v0.5.3

Table-column fidelity: FAILS. On the OCR fidelity check this build shifts config-table values by one column (every label carries the next column's value) while emitting well-formed output. Prefer level 5 or higher when table structure matters.

Vision weights are kept fp16; only the language tower is quantized.

Use

from mlx_vlm import load, generate
model, processor = load("1qh/chandra-ocr-2-oq4-mlx")
Downloads last month
27
Safetensors
Model size
1B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 1qh/chandra-ocr-2-oq4-mlx

Quantized
(35)
this model