shellkeeper-0.6b-GGUF

These are GGUF quants of hizkifw/shellkeeper-0.6b, a context-aware guard for AI agents' shell commands. The model reads the user's request, the agent's session so far, and the proposed command, then predicts safe or unsafe as a single token. You read the logprobs of those two tokens to get P(unsafe).

The prompt format, the label policy and the model's limitations are documented on the main model card.

Files

All quants except Q8_0 use an importance matrix computed on 1,500 training prompts (imatrix.gguf, included). "embQ3" means the tied 152k-vocab embedding is stored at Q3_K instead of llama.cpp's default Q6_K. At these sizes the embedding is about 40% of the file, and the change costs no accuracy.

file size held-out test AUC red-team r1 AUC red-team r2 AUC decisions โ‰  BF16 CPU p50
(BF16 reference) 1198 MB 0.990 0.954 0.938 โ€“ 154 ms
shellkeeper-0.6b-Q8_0.gguf 639 MB 0.990 0.954 0.938 0.2% 193 ms
shellkeeper-0.6b-Q4_K_M.gguf (recommended) 397 MB 0.989 0.951 0.937 0.7% 106 ms
shellkeeper-0.6b-Q3_K_M-embQ3.gguf 286 MB 0.989 0.950 0.938 1.9% 115 ms
shellkeeper-0.6b-Q2_K-embQ3.gguf 235 MB 0.986 0.942 0.928 3.9% 128 ms

How these were measured:

  • n = 3,007 across five eval sets. "Decisions โ‰  BF16" means the verdict at threshold 0.5 differs from the BF16 GGUF.
  • The BF16 GGUF itself matches HF transformers (mean |ฮ”p| 0.001).
  • CPU latency: i5-13600K, 6 threads, single request.
  • On a GPU (Vulkan, RX 7900 XTX) every quant takes about 22 ms.

Choosing a file. Q4_K_M is the default: no measurable loss, and the fastest on CPU. Use Q3_K_M-embQ3 for the smallest file with no measurable loss, and Q2_K-embQ3 if size matters more than the last bit of accuracy. Going below about 230 MB (IQ2_XXS, IQ1, ternary, even with QAT) degrades quality noticeably, so those files aren't published.

Usage with llama-server

llama-server -m shellkeeper-0.6b-Q4_K_M.gguf -c 4096 --port 8080
import math, requests

URL = "http://127.0.0.1:8080"
SAFE, UNSAFE = requests.post(f"{URL}/tokenize", json={"content": " safe unsafe"}).json()["tokens"]  # single tokens

def p_unsafe(prompt: str) -> float:
    r = requests.post(f"{URL}/completion", json={
        "prompt": prompt, "n_predict": 1, "n_probs": 50, "temperature": 0,
        "post_sampling_probs": False,   # raw model probabilities
        "samplers": [], "ignore_eos": True,
        # Equal bias on both labels forces a label to be *sampled* (the server withholds probs for byte-fragment
        # tokens). The reported pre-sampling probs and the safe/unsafe ratio are unaffected.
        "logit_bias": [[SAFE, 100], [UNSAFE, 100]],
    }).json()
    top = {t["id"]: t["logprob"] for t in r["completion_probabilities"][0]["top_logprobs"]}
    ls, lu = top.get(SAFE, -50.0), top.get(UNSAFE, -50.0)
    return 1 / (1 + math.exp(ls - lu))

prompt = """### Task
- clean and rebuild
### Session
shell: bash | cwd: /home/me/webapp
$ ls
build  node_modules  package.json  src  tests
### Command
rm -rf ./src
### Verdict:"""
print(p_unsafe(prompt))   # high; the same session with `rm -rf ./build` scores low
  • Use the exact prompt format from the main model card. Its Python build_prompt produces it.
  • Don't apply a chat template.
  • No BOS token is added (Qwen3 default).
  • In an agent loop, everything before ### Command changes little between calls. Enable cache_prompt so only the new command needs a prefill.

Policy

Treat the output as a score, not a verdict:

  • run automatically when p < lo
  • ask a human when lo โ‰ค p < hi
  • block when p โ‰ฅ hi

Tune lo and hi on your own traffic. Treat this as one layer next to sandboxing and least privilege, not the only security boundary.

Reproduce

These files were made with llama.cpp b11322:

python convert_hf_to_gguf.py shellkeeper-0.6b --outtype bf16 --outfile shellkeeper-0.6b-BF16.gguf
llama-quantize --imatrix imatrix.gguf shellkeeper-0.6b-BF16.gguf shellkeeper-0.6b-Q4_K_M.gguf Q4_K_M
llama-quantize --imatrix imatrix.gguf --token-embedding-type q3_K shellkeeper-0.6b-BF16.gguf shellkeeper-0.6b-Q3_K_M-embQ3.gguf Q3_K_M
llama-quantize --imatrix imatrix.gguf --token-embedding-type q3_K shellkeeper-0.6b-BF16.gguf shellkeeper-0.6b-Q2_K-embQ3.gguf Q2_K
llama-quantize shellkeeper-0.6b-BF16.gguf shellkeeper-0.6b-Q8_0.gguf Q8_0

The full quantization study, including QAT experiments, is in bench/QUANT.md in the GitHub repo.

SHA256

23bb7eaad429786ef2ea9e3140c16688e726bf7d01888d0b80ca28214b76612b  shellkeeper-0.6b-Q8_0.gguf
884482e9981217cae7235e8d0ba3b63b5a691462160cf1efec54d19b9b84f97f  shellkeeper-0.6b-Q4_K_M.gguf
976bbd8cbd96c64b9bd56e4d64e8df132c15c29c9a0b6f345f783ce8254f88c5  shellkeeper-0.6b-Q3_K_M-embQ3.gguf
ee1bc52b733c1c8acf011c3de73491a2e3edb714b4620bca58605f3131a280e7  shellkeeper-0.6b-Q2_K-embQ3.gguf
Downloads last month
147
GGUF
Model size
0.6B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for hizkifw/shellkeeper-0.6b-GGUF

Quantized
(1)
this model