Instructions to use Pitaya1219/Clef-Flash-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Pitaya1219/Clef-Flash-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Pitaya1219/Clef-Flash-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Pitaya1219/Clef-Flash-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Pitaya1219/Clef-Flash-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Pitaya1219/Clef-Flash-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Pitaya1219/Clef-Flash-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Pitaya1219/Clef-Flash-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Pitaya1219/Clef-Flash-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Pitaya1219/Clef-Flash-GGUF:Q4_K_M
Use Docker
docker model run hf.co/Pitaya1219/Clef-Flash-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use Pitaya1219/Clef-Flash-GGUF with Ollama:
ollama run hf.co/Pitaya1219/Clef-Flash-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use Pitaya1219/Clef-Flash-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Pitaya1219/Clef-Flash-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Pitaya1219/Clef-Flash-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Pitaya1219/Clef-Flash-GGUF with Docker Model Runner:
docker model run hf.co/Pitaya1219/Clef-Flash-GGUF:Q4_K_M
- Lemonade
How to use Pitaya1219/Clef-Flash-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Pitaya1219/Clef-Flash-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Clef-Flash-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Pitaya1219/Clef-Flash-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Pitaya1219/Clef-Flash-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Pitaya1219/Clef-Flash-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Pitaya1219/Clef-Flash-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Pitaya1219/Clef-Flash-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Pitaya1219/Clef-Flash-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Clef-Flash GGUF (imatrix Q4_K_M / IQ3 / IQ2)
Low-bit quantizations (Q4_K_M, IQ3_XXS, IQ2_M, IQ2_XS) of Cloudflare/clef-flash in the native clef GGUF architecture. Serve with llama-server b11415 or later (/v1/systemone, llama.cpp#29831).
Files
| file | head (dec.*, decision.*) |
notes |
|---|---|---|
Clef-Flash-Q4_K_M-imat.gguf |
q4_K | default llama-quantize Q4_K_M with the imatrix |
Clef-Flash-Q4_K_M-imat-head-q8.gguf |
q8_0 | same, with the joint schema head kept at Q8_0 |
Clef-Flash-IQ3_XXS-imat-head-q8.gguf |
q8_0 | 4.07 GB |
Clef-Flash-IQ2_M-imat-head-q8.gguf |
q8_0 | 3.74 GB |
Clef-Flash-IQ2_XS-imat-head-q8.gguf |
q8_0 | 3.42 GB |
How they were made
- Source: the BF16 GGUF from ggml-org/Clef-Flash-GGUF.
- imatrix:
clef-flash-imatrix.gguffrom Livesport/clef-flash-GGUF (Apache-2.0), computed on Clef-format decision records. All 248 backbone (blk.*) tensors matched; the head tensors have no imatrix data. llama-quantizefrom llama.cpp b11415; the head variants add--tensor-type "dec\..*=q8_0" --tensor-type "decision\..*=q8_0".
No text-generation or vision use. See the base model for the license and intended use.
Results
Run through llama-server b11415 (/v1/systemone, Vulkan, AMD Radeon 8060S), one request at a time, with the same 600 items for every model: 200 each from SST-2 (validation), AG News (test) and BoolQ (validation), sampled with a fixed seed. Agreement and KL are against Clef-Flash-Q8_0.gguf from ggml-org, not against bf16.
| model | size | SST-2 | AG News | BoolQ | mean | agreement | KL | BoolQ median latency |
|---|---|---|---|---|---|---|---|---|
| Q8_0 (ggml-org) | 9.66 GB | 97.5 | 89.0 | 93.0 | 93.2 | base | base | 272 ms |
| Q4_K_M (ggml-org) | 6.49 GB | 97.5 | 88.5 | 94.5 | 93.5 | 98.7 % | 0.0022 | 305 ms |
| Q4_K_M imat | 5.72 GB | 97.5 | 90.0 | 93.5 | 93.7 | 98.5 % | 0.0026 | 326 ms |
| Q4_K_M imat, head Q8_0 | 5.74 GB | 97.5 | 90.0 | 94.0 | 93.8 | 98.7 % | 0.0030 | 328 ms |
| IQ3_XXS imat, head Q8_0 | 4.07 GB | 96.0 | 89.0 | 93.5 | 92.8 | 98.0 % | 0.0122 | 348 ms |
| IQ2_M imat, head Q8_0 | 3.74 GB | 96.0 | 88.5 | 91.5 | 92.0 | 97.2 % | 0.0219 | 360 ms |
| IQ2_XS imat, head Q8_0 | 3.42 GB | 94.5 | 89.0 | 92.0 | 91.8 | 96.3 % | 0.0324 | 328 ms |
- Accuracy is the share of items where the most probable option is the gold label. With 200 items per task, a task score moves by ยฑ2-3 points from noise alone; agreement and KL are the steadier signals.
- Agreement and KL get worse monotonically as bits drop; accuracy falls about 1.4 points at IQ2_XS.
- Our imatrix builds are no better than ggml-org's Q4_K_M by KL, and keeping the head at Q8_0 did not help at Q4_K_M.
- Smaller files are not faster on Vulkan: the lower-bit files are 15-30 % slower than Q8_0.
- Only these three text tasks were measured, in English, with a single question per request; no vision input.
Size-constrained protection of IQ2_M
Livesport's per-tensor layout applied step by step to IQ2_M (all with the head at Q8_0), same 600 items, against the same Q8_0 baseline.
| model | size | mean acc. | agreement | KL |
|---|---|---|---|---|
| IQ2_M | 3.74 GB | 92.0 | 97.2 % | 0.0219 |
Clef-Flash-IQ2_M-gates.gguf (gated-delta-net gates F32) |
3.76 GB | 92.5 | 97.3 % | 0.0218 |
Clef-Flash-IQ2_M-gates-attn.gguf (+ attention Q8_0) |
4.09 GB | 93.2 | 97.3 % | 0.0213 |
Clef-Flash-IQ2_M-gates-attn-ssmout.gguf (+ ssm_out Q8_0) |
4.39 GB | 92.7 | 97.5 % | 0.0175 |
Clef-Flash-IQ2_M-dyn.gguf (+ token_embd Q8_0, early ffn_up/ffn_down Q6_K) |
5.31 GB | 93.2 | 98.0 % | 0.0124 |
| IQ3_XXS (head Q8_0) | 4.07 GB | 92.8 | 98.0 % | 0.0122 |
Protecting the gates and attention barely moves KL; ssm_out helps but costs 0.3 GB. At equal size a uniform IQ3_XXS is clearly better than any of the protected IQ2_M builds.
Raising the ffn tensors of IQ2_M to 3 bits
| model | size | mean acc. | agreement | KL |
|---|---|---|---|---|
| IQ2_M | 3.74 GB | 92.0 | 97.2 % | 0.0219 |
Clef-Flash-IQ2_M-ffn-down.gguf (ffn_down iq3_s) |
3.89 GB | 92.7 | 97.5 % | 0.0213 |
Clef-Flash-IQ2_M-ffn-updown.gguf (ffn_up + ffn_down iq3_s) |
4.07 GB | 93.0 | 97.8 % | 0.0191 |
Clef-Flash-IQ2_M-ffn-all.gguf (ffn_gate + ffn_up + ffn_down iq3_s) |
4.24 GB | 93.2 | 97.3 % | 0.0184 |
| IQ3_XXS (head Q8_0) | 4.07 GB | 92.8 | 98.0 % | 0.0122 |
Raising all ffn tensors costs 0.5 GB and only brings KL from 0.0219 to 0.0184. At the same size (4.07 GB) a uniform IQ3_XXS is about 1.6x closer to Q8_0 than ffn-updown. The error is spread across all tensors rather than concentrated in the ffn, so no targeted protection beat spending the bits uniformly.
Recipe files
recipe/imatrix-ja.gguf is the imatrix used for the *-imatja-* experiments (300 chunks of 512 tokens, computed on the Q8_0 with llama-imatrix -c 512 -b 512 -ub 512 -np 1 --parse-special). recipe/build_calib.py regenerates its calibration text: ~930 records in Clef's input format from public datasets (JMTEB massive/livedoor/amazon, WRIME, jawiki-paragraphs, plus English AG News/SST-2/BoolQ). llama-imatrix on a clef GGUF needs the batch size equal to the context size, because Clef accepts one sequence per batch.
- Downloads last month
- 9,921