Instructions to use MK4-Research/MK4-Ginko-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use MK4-Research/MK4-Ginko-v2 with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("MK4-Research/MK4-Ginko-v2") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use MK4-Research/MK4-Ginko-v2 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "MK4-Research/MK4-Ginko-v2"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "MK4-Research/MK4-Ginko-v2" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use MK4-Research/MK4-Ginko-v2 with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "MK4-Research/MK4-Ginko-v2"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "MK4-Research/MK4-Ginko-v2" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MK4-Research/MK4-Ginko-v2", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use MK4-Research/MK4-Ginko-v2 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "MK4-Research/MK4-Ginko-v2"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default MK4-Research/MK4-Ginko-v2
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use MK4-Research/MK4-Ginko-v2 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "MK4-Research/MK4-Ginko-v2"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "MK4-Research/MK4-Ginko-v2" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
MK4-Ginko v2
Ginko v2 is a standalone 6-bit MLX model for security code review. It traces the input to the sensitive operation, checks whether the displayed guard protects that exact value and path, and ends with VERDICT: SAFE or VERDICT: VULNERABLE <class>.
It scores 71.7% balanced accuracy on GuardBench, a held-out benchmark of 120 guard-reasoning snippets asked three ways. This release merges the selected LoRA into the base model in higher precision and then quantizes to 6-bit. Direct 4-bit fusion discarded enough of the LoRA update to change model behavior, so this repository uses 6-bit weights and needs no adapter at inference.
Training
The model begins from mlx-community/Qwen3.8-27B-4bit.
- A fresh v2 LoRA was trained on 181 v2 training examples for 300 steps: rank 32, 16 layers, learning rate
1e-5, 4-bit base, sequence length 512. - It then received a 75-step formatting and guard-path continuation on 91 curated examples: assistant-only loss, rank 32, 16 layers, learning rate
5e-6, sequence length 768. Every target ends in an explicit verdict line.
The continuation corpus passed the project's code-overlap check against held-out GuardBench and FBE items. It teaches short review output; it does not provide continuous memory compaction or persistent state between requests.
Evaluation
GuardBench contains 120 held-out snippets asked three ways (360 responses). Its headline is balanced accuracy: the mean of safe and vulnerable recall.
| Checkpoint | Balanced | Safe recall | Vulnerable recall | Missing verdicts |
|---|---|---|---|---|
| Prior v2 | 64.4% | 53/120 (44.2%) | 203/240 (84.6%) | 68/360 |
| v2 step 300 | 71.2% | 67/120 (55.8%) | 208/240 (86.7%) | 46/360 |
| Selected v2 adapter before merge |
73.3% | 63/120 (52.5%) | 226/240 (94.2%) | 0/360 |
| This repository (6-bit merge) | 71.7% | 57/120 (47.5%) | 230/240 (95.8%) | 0/360 |
GuardBench breakdown — this repository's 6-bit weights
| Guard relationship | Correct | Accuracy |
|---|---|---|
| No relevant guard | 56 / 60 | 93.3% |
| Guard covers the sensitive value | 26 / 60 | 43.3% |
| Guard covers it a second way | 31 / 60 | 51.7% |
| Guard checks the wrong value | 58 / 60 | 96.7% |
| Guard is irrelevant to the sink | 56 / 60 | 93.3% |
| Guard exists elsewhere, not on this path | 60 / 60 | 100.0% |
| Prompt form | Correct | Accuracy |
|---|---|---|
| Direct review prompt | 93 / 120 | 77.5% |
| Alternate wording | 97 / 120 | 80.8% |
| Structured review prompt | 97 / 120 | 80.8% |
Median completion length 50 tokens, deterministic decoding, no missing verdicts.
The two guard covers rows are the weak ones and they are the same question asked twice: a guard
is present and does protect the value. The model gets those right 43% and 52% of the time. Every
row where the correct answer is "vulnerable" scores above 93%. That asymmetry is the model's
defining characteristic, not a rounding artifact.
Additional benchmarks — measured on the adapter
| Benchmark | Score |
|---|---|
| MMLU-Pro (50-question local subset) | 27 / 50 (54%) |
| SecQA | 50 / 50 (100%) |
| CyberMetric | 47 / 50 (94%) |
| Local cyber multiple-choice | 50 / 50 (100%) |
| HumanEval (50-question local subset) | 42 / 50 (84%) |
| Vulnerability Analysis Benchmark (VAB) | 14 / 20 (70%) |
VAB category detail: broken-invariant 0/1, capability 1/1, multi-step chain 6/10, confidence calibration 4/4, missing-check 0/1, negative-space 1/1, root-cause 1/1, and state-sequence 1/1.
These suites were run against the adapter, not the 6-bit merge, and have not been repeated since. GuardBench was repeated on the merged weights and cost 1.6 points, so treat the numbers above as approximate for this repository rather than measured on it.
Every local model on the same benchmark
Twenty models were run through the identical protocol — 120 held-out snippets, three prompt phrasings each, 360 responses per model, deterministic decoding. This covers the whole project lineage plus the untrained bases as controls.
| Model | Balanced | Safe recall | Vulnerable recall | Missing verdicts |
|---|---|---|---|---|
| Ginko v2 selected adapter this model, before merge |
73.3% | 52.5% | 94.2% | 0/360 |
| Guard-fix pilot adapter earlier v2 experiment |
72.1% | 48.3% | 95.8% | 0/360 |
| Ginko v2 — the weights in this repository 6-bit merge |
71.7% | 47.5% | 95.8% | 0/360 |
| LOREA v5.9 35B MoE line |
71.5% | 49.2% | 93.8% | 10/360 |
| Ginko v2 step 300 before the formatting stage |
71.2% | 55.8% | 86.7% | 46/360 |
| Ginko v2.1 redo unreleased |
70.0% | 44.2% | 95.8% | 1/360 |
| Ginko v2.1 unreleased |
70.0% | 46.7% | 93.3% | 1/360 |
| Ginko v2 first run superseded |
64.4% | 44.2% | 84.6% | 68/360 |
| Ginko v1 previous release |
62.1% | 46.7% | 77.5% | 104/360 |
| LOREA v5.8 35B MoE line |
56.5% | 39.2% | 73.8% | 121/360 |
| LOREA v5.5 9B line |
52.5% | 10.8% | 94.2% | 10/360 |
| LOREA v6 Pilot 27B dense line |
42.9% | 31.7% | 54.2% | 183/360 |
| LOREA v5.7 9B line, final checkpoint |
40.6% | 25.8% | 55.4% | 142/360 |
| Qwythos 9B base untrained control |
38.1% | 15.0% | 61.3% | 142/360 |
| LOREA v5.4 9B line |
30.8% | 12.5% | 49.2% | 167/360 |
| LOREA v4.2 earliest local release |
30.8% | 11.7% | 50.0% | 177/360 |
| Qwen3.8-27B base untrained control — this model's base |
28.7% | 14.2% | 43.3% | 237/360 |
| Qwen3.6-27B base untrained control |
25.2% | 12.5% | 37.9% | 252/360 |
| LOREA v5.6 9B line, final checkpoint |
23.8% | 8.3% | 39.2% | 198/360 |
| 35B MoE base untrained control |
15.8% | 6.7% | 25.0% | 292/360 |
Three things in this table are worth stating plainly.
Every headline figure on this card is the 6-bit merge. The adapter scored 73.3% before merging; quantizing to 6 bit costs about 1.6 points, roughly six items. Both rows are shown so the cost is visible, but 71.7% is what this repository serves.
The top five models are within five items of each other. 73.3, 72.1, 71.7, 71.5 and 71.2 on a 360-response benchmark is not a ranking; it is a tie. One of those five is LOREA v5.9, trained earlier on a different 35B base. Treat "best model" claims across that group with suspicion.
Safe recall has not moved much. Sorted by the safe column instead, the best result in the whole project is 55.8% and this model is at 47.5%. What the later versions clearly did improve is vulnerable recall (77.5% to 95.8%) and output format reliability (104 missing verdicts to 0). Both matter, especially the second for anything parsing output. But roughly half of correctly-guarded code is still flagged, and that is the honest limit of this model.
The untrained bases score 15.8% to 28.7%, so the fine-tuning is doing real work — it has just been doing most of it on the easier half of the problem.
Limits
GuardBench still finds false alarms in guarded code: 63 of 120 safe prompts were marked vulnerable by these weights. Treat a finding as a review lead and validate it against the code. This model reviews focused snippets and does not provide complete codebase coverage.
Use
python3 -m mlx_lm generate --model MK4-Research/MK4-Ginko-v2 \
--prompt 'Review this code for security problems: <your code>'
The model is intended for defensive code review and research. Important security conclusions require human verification.
- Downloads last month
- 491
6-bit
