Clef-Flash GGUF (imatrix Q4_K_M / IQ3 / IQ2)

Low-bit quantizations (Q4_K_M, IQ3_XXS, IQ2_M, IQ2_XS) of Cloudflare/clef-flash in the native clef GGUF architecture. Serve with llama-server b11415 or later (/v1/systemone, llama.cpp#29831).

Files

file head (dec.*, decision.*) notes
Clef-Flash-Q4_K_M-imat.gguf q4_K default llama-quantize Q4_K_M with the imatrix
Clef-Flash-Q4_K_M-imat-head-q8.gguf q8_0 same, with the joint schema head kept at Q8_0
Clef-Flash-IQ3_XXS-imat-head-q8.gguf q8_0 4.07 GB
Clef-Flash-IQ2_M-imat-head-q8.gguf q8_0 3.74 GB
Clef-Flash-IQ2_XS-imat-head-q8.gguf q8_0 3.42 GB

How they were made

  • Source: the BF16 GGUF from ggml-org/Clef-Flash-GGUF.
  • imatrix: clef-flash-imatrix.gguf from Livesport/clef-flash-GGUF (Apache-2.0), computed on Clef-format decision records. All 248 backbone (blk.*) tensors matched; the head tensors have no imatrix data.
  • llama-quantize from llama.cpp b11415; the head variants add --tensor-type "dec\..*=q8_0" --tensor-type "decision\..*=q8_0".

No text-generation or vision use. See the base model for the license and intended use.

Results

Run through llama-server b11415 (/v1/systemone, Vulkan, AMD Radeon 8060S), one request at a time, with the same 600 items for every model: 200 each from SST-2 (validation), AG News (test) and BoolQ (validation), sampled with a fixed seed. Agreement and KL are against Clef-Flash-Q8_0.gguf from ggml-org, not against bf16.

model size SST-2 AG News BoolQ mean agreement KL BoolQ median latency
Q8_0 (ggml-org) 9.66 GB 97.5 89.0 93.0 93.2 base base 272 ms
Q4_K_M (ggml-org) 6.49 GB 97.5 88.5 94.5 93.5 98.7 % 0.0022 305 ms
Q4_K_M imat 5.72 GB 97.5 90.0 93.5 93.7 98.5 % 0.0026 326 ms
Q4_K_M imat, head Q8_0 5.74 GB 97.5 90.0 94.0 93.8 98.7 % 0.0030 328 ms
IQ3_XXS imat, head Q8_0 4.07 GB 96.0 89.0 93.5 92.8 98.0 % 0.0122 348 ms
IQ2_M imat, head Q8_0 3.74 GB 96.0 88.5 91.5 92.0 97.2 % 0.0219 360 ms
IQ2_XS imat, head Q8_0 3.42 GB 94.5 89.0 92.0 91.8 96.3 % 0.0324 328 ms
  • Accuracy is the share of items where the most probable option is the gold label. With 200 items per task, a task score moves by ยฑ2-3 points from noise alone; agreement and KL are the steadier signals.
  • Agreement and KL get worse monotonically as bits drop; accuracy falls about 1.4 points at IQ2_XS.
  • Our imatrix builds are no better than ggml-org's Q4_K_M by KL, and keeping the head at Q8_0 did not help at Q4_K_M.
  • Smaller files are not faster on Vulkan: the lower-bit files are 15-30 % slower than Q8_0.
  • Only these three text tasks were measured, in English, with a single question per request; no vision input.

Size-constrained protection of IQ2_M

Livesport's per-tensor layout applied step by step to IQ2_M (all with the head at Q8_0), same 600 items, against the same Q8_0 baseline.

model size mean acc. agreement KL
IQ2_M 3.74 GB 92.0 97.2 % 0.0219
Clef-Flash-IQ2_M-gates.gguf (gated-delta-net gates F32) 3.76 GB 92.5 97.3 % 0.0218
Clef-Flash-IQ2_M-gates-attn.gguf (+ attention Q8_0) 4.09 GB 93.2 97.3 % 0.0213
Clef-Flash-IQ2_M-gates-attn-ssmout.gguf (+ ssm_out Q8_0) 4.39 GB 92.7 97.5 % 0.0175
Clef-Flash-IQ2_M-dyn.gguf (+ token_embd Q8_0, early ffn_up/ffn_down Q6_K) 5.31 GB 93.2 98.0 % 0.0124
IQ3_XXS (head Q8_0) 4.07 GB 92.8 98.0 % 0.0122

Protecting the gates and attention barely moves KL; ssm_out helps but costs 0.3 GB. At equal size a uniform IQ3_XXS is clearly better than any of the protected IQ2_M builds.

Raising the ffn tensors of IQ2_M to 3 bits

model size mean acc. agreement KL
IQ2_M 3.74 GB 92.0 97.2 % 0.0219
Clef-Flash-IQ2_M-ffn-down.gguf (ffn_down iq3_s) 3.89 GB 92.7 97.5 % 0.0213
Clef-Flash-IQ2_M-ffn-updown.gguf (ffn_up + ffn_down iq3_s) 4.07 GB 93.0 97.8 % 0.0191
Clef-Flash-IQ2_M-ffn-all.gguf (ffn_gate + ffn_up + ffn_down iq3_s) 4.24 GB 93.2 97.3 % 0.0184
IQ3_XXS (head Q8_0) 4.07 GB 92.8 98.0 % 0.0122

Raising all ffn tensors costs 0.5 GB and only brings KL from 0.0219 to 0.0184. At the same size (4.07 GB) a uniform IQ3_XXS is about 1.6x closer to Q8_0 than ffn-updown. The error is spread across all tensors rather than concentrated in the ffn, so no targeted protection beat spending the bits uniformly.

Recipe files

recipe/imatrix-ja.gguf is the imatrix used for the *-imatja-* experiments (300 chunks of 512 tokens, computed on the Q8_0 with llama-imatrix -c 512 -b 512 -ub 512 -np 1 --parse-special). recipe/build_calib.py regenerates its calibration text: ~930 records in Clef's input format from public datasets (JMTEB massive/livedoor/amazon, WRIME, jawiki-paragraphs, plus English AG News/SST-2/BoolQ). llama-imatrix on a clef GGUF needs the batch size equal to the context size, because Clef accepts one sequence per batch.

Downloads last month
9,921
GGUF
Model size
9B params
Architecture
clef
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Pitaya1219/Clef-Flash-GGUF

Finetuned
Qwen/Qwen3.5-9B
Quantized
(53)
this model

Space using Pitaya1219/Clef-Flash-GGUF 1