Instructions to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF # Run inference directly in the terminal: llama cli -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF # Run inference directly in the terminal: llama cli -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF # Run inference directly in the terminal: ./llama-cli -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF
Use Docker
docker model run hf.co/IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF
- LM Studio
- Jan
- Ollama
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF with Ollama:
ollama run hf.co/IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF
- Unsloth Desktop
- Pi
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF with Docker Model Runner:
docker model run hf.co/IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF
- Lemonade
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF
Run and chat with the model
lemonade run user.Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Qwen3.8-Flash-Coder APEX-I-MiniPlus-V2.1 GGUF
- Optimization History & Transparency Notice
- Model Files & Technical Specifications
- Surgical Tensor Quantization Map (Audited from GGUF)
- Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50 & RAM Streaming)
- The 24GB Miracle: Full 256K Context Runs In VRAM!
- Recommended Configuration & Setup
- 7. Community Compute Fund & Priority Model Requests
- Optimization History & Transparency Notice
Quick Navigation Index
Independently computed on private cloud clusters. If this handcrafted release saves you VRAM and runs faster on your GPU, consider fueling the community compute fund on Ko-fi.
- Optimization History & Transparency Notice
- Empirical Benchmarks & Fidelity Verification
- Quality Spectrum: APEX-I-MiniPlus V2.1 vs. Standard Flat Quantizations
- Model Files & Technical Specifications
- Surgical Tensor Quantization Map (Audited from GGUF)
- Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50 & RAM Streaming)
- The 24GB Miracle: Full 256K Context Runs In VRAM!
- Recommended Configuration & Setup
- Recommended Generation Parameters
- CRITICAL: Coding Syntax & Repeat Penalty Advisory (Preventing Character Swapping)
- Hardened Agentic Chat Template & Reasoning Effort
- Optional Support
EXPERIMENTAL PRE-RELEASE NOTICE: ENGLISH-ONLY CODING SPECIALIST
This model suite is quantized from Jab1718/qwen3.8-flash-coder-85gb-bf16, which is an intermediate experimental slice created using
moe-slice(352 out of 512 routed experts were permanently pruned exclusively against English Python and SWE-bench calibration datasets).
- English Coding Only: This model is strictly designed for programming, code completion, refactoring, and agentic tool-calling in English.
- Severe Multilingual & General Degradation: Because conversational and multilingual experts were pruned and the upstream author has not yet released the recovery fine-tuning pass, this model severely degrades and outputs broken text in languages other than English (e.g., Spanish, French, German, etc.) or in general chit-chat.
- Incompatible with Strata Engine: This model uses a 160-expert layout with decoupled n-gram tables; it is not compatible with Strata Engine (which requires the 512-expert monolith and 51B PLE tables). Run using stock
llama.cpp(llama-server) or LM Studio.
Qwen3.8-Flash-Coder APEX-I-MiniPlus-V2.1 GGUF
The Definitive Frontier MoE · Efficient System RAM Offload · Full 256K Context on 24GB Workstations
EXPLORE THE COMPLETE QWEN3.8 FLASH CODER LINEUP
These are complementary APEX-I releases, not alternate downloads of the same model:
- Qwen3.8-Flash-Coder APEX-I-MiniPlus-V2.1 - specialist agentic coding MoE tuned for Q5-Q6 quality (21.77 GB / 3.45 BPW).
- Qwen3.8-Flash-Coder APEX-I-NanoPlus - ultra-compact footprint achieving solid Q4 quality (18.34 GB / 2.90 BPW).
- Qwen3.8-Flash-Coder-85GB Lossless BF16 GGUF - uncompressed reference baseline (85.30 GB / 16.00 BPW).
THE DEFINITIVE SPECIFICATION IN THE 21–22 GB CEILING
This APEX-I-MiniPlus-V2.1 release represents the specialized tensor-by-tensor configuration for sparse Mixture-of-Experts quantization within a 21–22 GB envelope. Every tensor across its 48 hybrid layers (36 linear SSM DeltaNet + 12 sparse full attention) and 160 MoE experts has been mathematically allocated to maximize reasoning precision, preserve routing behavior, and prevent avoidable CPU dequantization stalls.
EMPIRICAL BENCHMARK & QUALITY COMPARISON
Quantization Specification File Size (Disk) Memory Footprint (RAM/VRAM) Average BPW WikiText-2 Perplexity Delta PPL vs BF16 (%) Quality Tier Equivalent Uncompressed BF16 Reference 85.30 GB(79.44 GiB)79.44 GiB16.00 BPW 30.0975 +/- 0.1200Baseline (0.00%) Lossless Reference Baseline APEX-I-MiniPlus V2.1 (CURRENT) 21.77 GB(20.27 GiB)20.27 GiB3.45 BPW 30.1495 +/- 1.0089+0.0520 (+0.17%)Q5_K_L / Q6_K Tier Boundary APEX-I-NanoPlus 18.34 GB(17.08 GiB)17.08 GiB2.90 BPW 34.4199 +/- 1.1591+4.3224 (+14.36%)Solid Q4_K_M Tier Standard Flat Q4_K_M 28.39 GB 26.44 GiB 4.50 BPW aprox. 30.28 - 30.38 +0.18 a +0.28 (+0.7%) Standard industry trade-off Standard Flat Q3_K_M 22.22 GB 20.69 GiB 3.44 BPW aprox. 30.55 - 30.95 +0.45 a +0.85 (+2.1%) Noticeable syntax drop & bracket noise Generic APEX Mini (IQ2_S) 17.73 GB 16.51 GiB 2.50 BPW aprox. 31.60 - 33.10+ +1.50 a +3.00+ (+7.5%) Severe reasoning breakdown Routing Fidelity: All router gates (
ffn_gate_inp,ffn_gate_inp_shexp) remain in uncompressedF32, preserving zero routing drift across 160 experts.
- Q6_K-Bordering Tier in Language Fidelity: WikiText-2 perplexity delta is exceptionally low (+0.0520 / +0.17% vs. BF16 baseline), placing overall code representation at the boundary of a 38 GB
Q6_Kbuild within an agile 21.77 GB footprint.- High-Precision Foundation Knowledge: Shared foundation experts run in uncompressed
Q8_0(down, block-32) +Q6_K(gate/up) across all 48 layers, keeping coding syntax intact across 100% of tokens.- Bypassed PLE Overhead: The 51B parameter N-gram table was cleanly removed (
ple_layer_ids: []), enabling 100% GPU VRAM execution with zero host RAM overhead.
DO NOT CONFUSE APEX-I-MINIPLUS WITH GENERIC COMMUNITY APEX-I-MINI!
- Generic Community APEX-I-Mini: Uniformly compresses all core MoE experts down to aggressive 2-bit
IQ2_S, leaves the sensitive token output head unarmored at 3-bitQ3_K_M, and compresses attention projections down toQ3_K. In deep reasoning models, this triggers severe perplexity spikes, syntax errors, and broken code brackets.- Handcrafted APEX-I-MiniPlus V2.1: Preserves specified router gates in uncompressed
F32, armors the token output head in high-precisionQ6_K, safeguards attention gates inQ8_0, and keeps core reasoning experts at calibrated 3-bit to 4-bit non-linear quantization.
Optimization History & Transparency Notice
We maintain our releases publicly as a transparent engineering record of continuous optimization. Below is the exact evolutionary roadmap of our MiniPlus architectures:
| Specification | Core Experts (10 Active/160) | Shared Expert (shexp) |
Sparse Attention (12 Layers) | Highway Connections (4-Way) | Attention Gates (36 Layers) | Output Head (output.weight) |
Routers (gate_inp) |
Real-World Impact |
|---|---|---|---|---|---|---|---|---|
| Generic APEX Mini | IQ2_S (2.50 bpw) |
Q4_K / Q3_K |
Q3_K |
Compressed | Compressed | Q3_K_M |
Compressed | High perplexity spikes, syntax errors in code generation. |
| MiniPlus V2.1 (CURRENT) | IQ3_XXS & IQ4_NL |
Q8_0 + Q6_K (All 48 layers) |
Q4_K (q/k/v) + Q6_K (output) |
Q8_0 |
Q8_0 |
Q6_K |
F32 |
Definitive Build (21.77 GB). Zero AVX2 CPU stalls, zero divisibility crashes, and full 256K context support. |
Model Files & Technical Specifications
| File Name | File Size | Memory Footprint | BPW | Description |
|---|---|---|---|---|
Qwen3.8-Flash-Coder-85GB.APEX-I-MiniPlus-V2.1.gguf |
21.77 GB (20.27 GiB) |
20.27 GiB |
3.45 BPW | Frontier agentic coding, SWE-bench vulnerability analysis, reasoning & hybrid SSM/sparse MoE |
- Base Model: Jab1718/qwen3.8-flash-coder-85gb-bf16 (derived from
Qwen/Qwen3.8-Flash-Nextsliced withmoe-sliceand DoRA calibration) - Parameters: 45.8B total (approx. 3.7B active per token: 10 routed experts + 1 shared expert + dense backbone; 4.9B with vocabulary embeddings)
- Architecture:
Qwen4ExpForCausalLM(48 hybrid layers: 36 linear SSM DeltaNet + 12 sparse full attention, 160 MoE experts per layer with 10 active) - Context Length: 262,144 tokens (native 256K)
- PLE / N-Gram Table: Decoupled and bypassed (
ple_layer_ids: []) for 100% GPU VRAM execution.
Surgical Tensor Quantization Map (Audited from GGUF)
The exact tensor breakdown below has been verified directly from the compiled binary weights:
| Layer Group | Sub-Component / Tensor | Qty | Precision | Engineering Rationale |
|---|---|---|---|---|
| Global Output Head | output.weight |
1 | Q6_K |
Preserves near-FP16 token classification; eliminates syntax errors, bracket drops, and hallucinations. |
| Global Embeddings | token_embd.weight |
1 | Q4_K |
High-fidelity vocabulary embedding representation across 248k vocabulary. |
| All Normalizations | output_norm, attn_norm, ffn_norm, hc_norm, ssm_norm |
146 | F32 |
100% uncompressed numerical stability across all 48 layers. |
| Expert Routers | blk.*.ffn_gate_inp, ffn_gate_inp_shexp |
96 | F32 |
100% uncompressed routing fidelity across 160 experts; zero router drift. |
| Attention Gates | blk.*.attn_gate.weight (36 SSM Layers) |
36 | Q8_0 |
High-precision attention gating across hybrid DeltaNet recurrence layers; eliminates crosstalk. |
| Linear Attention Projections | blk.*.attn_qkv.weight (36 SSM Layers) |
36 | Q4_K |
Balanced precision for input state-space projections. |
| SSM Linear Output | blk.*.ssm_out.weight (36 SSM Layers) |
36 | Q6_K |
Armored state-space output projection; preserves recurrence channel mixing. |
| Recurrent SSM Parameters | blk.*.ssm_{a,alpha,beta,conv1d,dt,norm} |
216 | F32 |
Guarded in uncompressed FP32 to prevent DeltaNet recurrent state drift. |
| Periodic Sparse Attention | blk.{3,7,...}.attn_{q,k,v}.weight (12 Layers) |
36 | Q4_K |
Quadratic full-attention anchor checkpoints for deep needle-in-a-haystack retrieval. |
| Periodic Sparse Attention Output | blk.{3,7,...}.attn_output.weight (12 Layers) |
12 | Q6_K |
High-precision attention output projection over deep context. |
| QSA Sparse Attention Indexer | blk.{3,7,...}.indexer.{q,k}_proj.weight |
24 | BF16 |
High-fidelity sparse indexer projections for query-stream attention routing. |
| QSA Indexer Norms | blk.{3,7,...}.indexer.{q,k}_norm.weight |
24 | F32 |
Uncompressed indexer layer normalizations. |
| Shared Foundation Experts | blk.*.ffn_down_shexp.weight (All 48 Layers) |
48 | Q8_0 |
ne0=640 in standard block-32; eliminates divisibility validation aborts while preserving foundation knowledge. |
| Shared Foundation Experts | blk.*.ffn_{gate,up}_shexp.weight (All 48 Layers) |
96 | Q6_K |
High-precision shared backbone active on 100% of tokens. |
| Highway Connections (4-Way) | blk.*.hc_{attn,ffn}_{down,inject,up}.weight |
288 | Q8_0 |
High-speed residual highway bypass; ne0=320 up-projections mapped to block-32 Q8_0. |
| Routed MoE Down-Projections | blk.*.ffn_down_exps.weight (All 48 Layers) |
48 | IQ4_NL (36) / Q4_0 (12) |
Non-linear codebook quantization for SSM layers, Q4_0 for anchor layers. |
| Routed MoE Gate/Up Projections | blk.*.ffn_{gate,up}_exps.weight (All 48 Layers) |
96 | IQ3_XXS (72) / Q4_K (24) |
Calibrated with imatrix for maximum compactness, with Q4_K protection on anchor layers. |
| Residual Output Highway | output_hc_{down,up}.weight |
2 | Q4_0 |
Low-rank residual highway projections at model termination. |
| Residual Output Highway Norm | output_hc_norm.weight |
1 | F32 |
Final residual normalization anchor. |
Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50 & RAM Streaming)
| Hardware Target | Offload Mode | Generation Speed (Est.) | Prompt Prefill Speed (Est.) | Highlights |
|---|---|---|---|---|
| NVIDIA RTX 5080 / 5090 (Blackwell) | Full GPU (-ngl 99) |
approx. 210 – 240 tok/s | 2,600 – 3,500+ tok/s | Blistering throughput on Hopper / Blackwell architectures |
| NVIDIA RTX 4090 (24GB GDDR6X) | Full GPU (-ngl 99) |
80 – 105+ tok/s | 1,800 – 2,400+ tok/s | Hybrid linear attention slashes prefill latency |
| NVIDIA RTX 3090 (24GB GDDR6) | Full GPU (-ngl 99) |
65 – 82+ tok/s | 1,300 – 1,800+ tok/s | Full 256k native window in VRAM |
| Workstation / Laptop (DDR4 / DDR5 RAM) | Hybrid Offload (Few layers in VRAM) | Hardware-dependent | Hardware-dependent | Zero AVX2 CPU stalls; efficient streaming from system RAM |
The 24GB Miracle: Full 256K Context Runs In VRAM!
Qwen3.8-Flash-Coder-85GB APEX-I-MiniPlus-V2.1 fits deep context windows within standard 24GB and 32GB GPUs:
| Context Length | Model Weights | KV Cache (q8_0) | Compute Buffers | Total GPU VRAM (Est.) | Feasibility |
|---|---|---|---|---|---|
| 32,768 (32k) | 20.27 GiB |
0.45 GiB |
1.50 GiB |
22.22 GiB |
Effortless fit on 24GB GPUs (RTX 3090 / 4090) |
| 65,536 (64k) | 20.27 GiB |
0.85 GiB |
1.65 GiB |
22.77 GiB |
Full offload on 24GB GPUs |
| 131,072 (128k) | 20.27 GiB |
1.45 GiB |
1.90 GiB |
23.62 GiB |
Full offload on 24GB GPUs |
| 262,144 (256k) | 20.27 GiB |
2.60 GiB |
2.40 GiB |
25.27 GiB |
Full offload on 32GB (RTX 5090) or partial RAM offload |
Recommended Configuration & Setup
llama-server \
-m Qwen3.8-Flash-Coder-85GB.APEX-I-MiniPlus-V2.1.gguf \
--jinja \
-ngl 99 \
--ctx-size 65536 \
--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.00 \
--port 8080
Recommended Generation Parameters
| Hyperparameter | Value | Description / Creator Notice |
|---|---|---|
| Temperature | 0.60 |
Recommended default for code generation & reasoning (--temp 0.6). |
| Top-P | 0.95 |
Source sampling parameter (--top-p 0.95). |
| Top-K | 20 |
Source sampling parameter (--top-k 20). |
| Min-P | 0.00 |
Source sampling parameter (--min-p 0.00). |
| Template Engine | --jinja |
Recommended official chat template flag. |
| Context Size | 65536 |
High-throughput 64K context (up to native 256K). |
CRITICAL ADVISORY FOR CODING WORKFLOWS: PREVENTING SYNTAX & TOKEN SWAPPING
In programming code, brackets (
{,}), assignment operators (=), and indentation whitespace repeat constantly across multi-line structures.Common Issue: Many local frontends (such as LM Studio defaults, Ollama, or web interfaces) ship with
repeat_penaltyset to1.1or1.15. While this prevents loops in creative prose, applying repeat penalties to code artificially penalizes necessary syntax tokens. When the logit of{drops, the model is forced to emit the next closest mathematical token (=or[), resulting in character swapping or dropped/doubled whitespace.Eliminating Character Swapping:
- Disable Repeat Penalties (Required for Code):
repeat_penalty: 1.0(strictly disabled)presence_penalty: 0.0frequency_penalty: 0.0- Calibrate Samplers:
temperature: 0.60(or0.20-0.30for strict, deterministic code syntax)min_p: 0.05(prunes low-probability noise tokens effectively)top_p: 0.95top_k: 20- Native Jinja Formatting: Always pass the
--jinjaflag so the tokenizer handles leading-space BPE tokens cleanly.
HARDENED AGENTIC CHAT TEMPLATE (JINJA)
The repository supports full Jinja chat templating with multi-level reasoning effort control:
low/minimal: Keeps internal thinking concise and focused strictly on immediate execution steps to minimize latency in automated loops.medium(default): Balanced, structured reasoning process with standard analytical depth.high/xhigh: Guides the model to formulate a clear implementation plan upfront before generating code.
7. Community Compute Fund & Priority Model Requests
All IsValorum quantizations will always remain completely free and open to the public without paywalls.
However, cloud GPU compute is expensive. If you find these builds valuable and would like to support the project or request a specific model architecture to be prioritized for the next MiniPlus/NanoPlus release, you can sponsor GPU compute time through Ko-fi:
(When supporting on Ko-fi, feel free to leave a note with your Hugging Face handle and the specific model you would like prioritized).
- Downloads last month
- 1,804
We're not able to determine the quantization variants.
Model tree for IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF
Base model
Qwen/Qwen3.8-Flash-Next