Quick Navigation Index

Support on Ko-fi

Independently computed on private cloud clusters. If this handcrafted release saves you VRAM and runs faster on your GPU, consider fueling the community compute fund on Ko-fi.

  1. Optimization History & Transparency Notice
  2. Empirical Benchmarks & Fidelity Verification
  3. Quality Spectrum: APEX-I-MiniPlus V2.1 vs. Standard Flat Quantizations
  4. Model Files & Technical Specifications
  5. Surgical Tensor Quantization Map (Audited from GGUF)
  6. Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50 & RAM Streaming)
  7. The 24GB Miracle: Full 256K Context Runs In VRAM!
  8. Recommended Configuration & Setup
  9. Recommended Generation Parameters
  10. CRITICAL: Coding Syntax & Repeat Penalty Advisory (Preventing Character Swapping)
  11. Hardened Agentic Chat Template & Reasoning Effort
  12. Optional Support

EXPERIMENTAL PRE-RELEASE NOTICE: ENGLISH-ONLY CODING SPECIALIST

This model suite is quantized from Jab1718/qwen3.8-flash-coder-85gb-bf16, which is an intermediate experimental slice created using moe-slice (352 out of 512 routed experts were permanently pruned exclusively against English Python and SWE-bench calibration datasets).

  • English Coding Only: This model is strictly designed for programming, code completion, refactoring, and agentic tool-calling in English.
  • Severe Multilingual & General Degradation: Because conversational and multilingual experts were pruned and the upstream author has not yet released the recovery fine-tuning pass, this model severely degrades and outputs broken text in languages other than English (e.g., Spanish, French, German, etc.) or in general chit-chat.
  • Incompatible with Strata Engine: This model uses a 160-expert layout with decoupled n-gram tables; it is not compatible with Strata Engine (which requires the 512-expert monolith and 51B PLE tables). Run using stock llama.cpp (llama-server) or LM Studio.

Qwen3.8-Flash-Coder APEX-I-MiniPlus-V2.1 GGUF

The Definitive Frontier MoE · Efficient System RAM Offload · Full 256K Context on 24GB Workstations

EXPLORE THE COMPLETE QWEN3.8 FLASH CODER LINEUP

These are complementary APEX-I releases, not alternate downloads of the same model:

THE DEFINITIVE SPECIFICATION IN THE 21–22 GB CEILING

This APEX-I-MiniPlus-V2.1 release represents the specialized tensor-by-tensor configuration for sparse Mixture-of-Experts quantization within a 21–22 GB envelope. Every tensor across its 48 hybrid layers (36 linear SSM DeltaNet + 12 sparse full attention) and 160 MoE experts has been mathematically allocated to maximize reasoning precision, preserve routing behavior, and prevent avoidable CPU dequantization stalls.

EMPIRICAL BENCHMARK & QUALITY COMPARISON

Quantization Specification File Size (Disk) Memory Footprint (RAM/VRAM) Average BPW WikiText-2 Perplexity Delta PPL vs BF16 (%) Quality Tier Equivalent
Uncompressed BF16 Reference 85.30 GB (79.44 GiB) 79.44 GiB 16.00 BPW 30.0975 +/- 0.1200 Baseline (0.00%) Lossless Reference Baseline
APEX-I-MiniPlus V2.1 (CURRENT) 21.77 GB (20.27 GiB) 20.27 GiB 3.45 BPW 30.1495 +/- 1.0089 +0.0520 (+0.17%) Q5_K_L / Q6_K Tier Boundary
APEX-I-NanoPlus 18.34 GB (17.08 GiB) 17.08 GiB 2.90 BPW 34.4199 +/- 1.1591 +4.3224 (+14.36%) Solid Q4_K_M Tier
Standard Flat Q4_K_M 28.39 GB 26.44 GiB 4.50 BPW aprox. 30.28 - 30.38 +0.18 a +0.28 (+0.7%) Standard industry trade-off
Standard Flat Q3_K_M 22.22 GB 20.69 GiB 3.44 BPW aprox. 30.55 - 30.95 +0.45 a +0.85 (+2.1%) Noticeable syntax drop & bracket noise
Generic APEX Mini (IQ2_S) 17.73 GB 16.51 GiB 2.50 BPW aprox. 31.60 - 33.10+ +1.50 a +3.00+ (+7.5%) Severe reasoning breakdown

Routing Fidelity: All router gates (ffn_gate_inp, ffn_gate_inp_shexp) remain in uncompressed F32, preserving zero routing drift across 160 experts.

  • Q6_K-Bordering Tier in Language Fidelity: WikiText-2 perplexity delta is exceptionally low (+0.0520 / +0.17% vs. BF16 baseline), placing overall code representation at the boundary of a 38 GB Q6_K build within an agile 21.77 GB footprint.
  • High-Precision Foundation Knowledge: Shared foundation experts run in uncompressed Q8_0 (down, block-32) + Q6_K (gate/up) across all 48 layers, keeping coding syntax intact across 100% of tokens.
  • Bypassed PLE Overhead: The 51B parameter N-gram table was cleanly removed (ple_layer_ids: []), enabling 100% GPU VRAM execution with zero host RAM overhead.

DO NOT CONFUSE APEX-I-MINIPLUS WITH GENERIC COMMUNITY APEX-I-MINI!

  • Generic Community APEX-I-Mini: Uniformly compresses all core MoE experts down to aggressive 2-bit IQ2_S, leaves the sensitive token output head unarmored at 3-bit Q3_K_M, and compresses attention projections down to Q3_K. In deep reasoning models, this triggers severe perplexity spikes, syntax errors, and broken code brackets.
  • Handcrafted APEX-I-MiniPlus V2.1: Preserves specified router gates in uncompressed F32, armors the token output head in high-precision Q6_K, safeguards attention gates in Q8_0, and keeps core reasoning experts at calibrated 3-bit to 4-bit non-linear quantization.

Optimization History & Transparency Notice

We maintain our releases publicly as a transparent engineering record of continuous optimization. Below is the exact evolutionary roadmap of our MiniPlus architectures:

Specification Core Experts (10 Active/160) Shared Expert (shexp) Sparse Attention (12 Layers) Highway Connections (4-Way) Attention Gates (36 Layers) Output Head (output.weight) Routers (gate_inp) Real-World Impact
Generic APEX Mini IQ2_S (2.50 bpw) Q4_K / Q3_K Q3_K Compressed Compressed Q3_K_M Compressed High perplexity spikes, syntax errors in code generation.
MiniPlus V2.1 (CURRENT) IQ3_XXS & IQ4_NL Q8_0 + Q6_K (All 48 layers) Q4_K (q/k/v) + Q6_K (output) Q8_0 Q8_0 Q6_K F32 Definitive Build (21.77 GB). Zero AVX2 CPU stalls, zero divisibility crashes, and full 256K context support.

Model Files & Technical Specifications

File Name File Size Memory Footprint BPW Description
Qwen3.8-Flash-Coder-85GB.APEX-I-MiniPlus-V2.1.gguf 21.77 GB (20.27 GiB) 20.27 GiB 3.45 BPW Frontier agentic coding, SWE-bench vulnerability analysis, reasoning & hybrid SSM/sparse MoE
  • Base Model: Jab1718/qwen3.8-flash-coder-85gb-bf16 (derived from Qwen/Qwen3.8-Flash-Next sliced with moe-slice and DoRA calibration)
  • Parameters: 45.8B total (approx. 3.7B active per token: 10 routed experts + 1 shared expert + dense backbone; 4.9B with vocabulary embeddings)
  • Architecture: Qwen4ExpForCausalLM (48 hybrid layers: 36 linear SSM DeltaNet + 12 sparse full attention, 160 MoE experts per layer with 10 active)
  • Context Length: 262,144 tokens (native 256K)
  • PLE / N-Gram Table: Decoupled and bypassed (ple_layer_ids: []) for 100% GPU VRAM execution.

Surgical Tensor Quantization Map (Audited from GGUF)

The exact tensor breakdown below has been verified directly from the compiled binary weights:

Layer Group Sub-Component / Tensor Qty Precision Engineering Rationale
Global Output Head output.weight 1 Q6_K Preserves near-FP16 token classification; eliminates syntax errors, bracket drops, and hallucinations.
Global Embeddings token_embd.weight 1 Q4_K High-fidelity vocabulary embedding representation across 248k vocabulary.
All Normalizations output_norm, attn_norm, ffn_norm, hc_norm, ssm_norm 146 F32 100% uncompressed numerical stability across all 48 layers.
Expert Routers blk.*.ffn_gate_inp, ffn_gate_inp_shexp 96 F32 100% uncompressed routing fidelity across 160 experts; zero router drift.
Attention Gates blk.*.attn_gate.weight (36 SSM Layers) 36 Q8_0 High-precision attention gating across hybrid DeltaNet recurrence layers; eliminates crosstalk.
Linear Attention Projections blk.*.attn_qkv.weight (36 SSM Layers) 36 Q4_K Balanced precision for input state-space projections.
SSM Linear Output blk.*.ssm_out.weight (36 SSM Layers) 36 Q6_K Armored state-space output projection; preserves recurrence channel mixing.
Recurrent SSM Parameters blk.*.ssm_{a,alpha,beta,conv1d,dt,norm} 216 F32 Guarded in uncompressed FP32 to prevent DeltaNet recurrent state drift.
Periodic Sparse Attention blk.{3,7,...}.attn_{q,k,v}.weight (12 Layers) 36 Q4_K Quadratic full-attention anchor checkpoints for deep needle-in-a-haystack retrieval.
Periodic Sparse Attention Output blk.{3,7,...}.attn_output.weight (12 Layers) 12 Q6_K High-precision attention output projection over deep context.
QSA Sparse Attention Indexer blk.{3,7,...}.indexer.{q,k}_proj.weight 24 BF16 High-fidelity sparse indexer projections for query-stream attention routing.
QSA Indexer Norms blk.{3,7,...}.indexer.{q,k}_norm.weight 24 F32 Uncompressed indexer layer normalizations.
Shared Foundation Experts blk.*.ffn_down_shexp.weight (All 48 Layers) 48 Q8_0 ne0=640 in standard block-32; eliminates divisibility validation aborts while preserving foundation knowledge.
Shared Foundation Experts blk.*.ffn_{gate,up}_shexp.weight (All 48 Layers) 96 Q6_K High-precision shared backbone active on 100% of tokens.
Highway Connections (4-Way) blk.*.hc_{attn,ffn}_{down,inject,up}.weight 288 Q8_0 High-speed residual highway bypass; ne0=320 up-projections mapped to block-32 Q8_0.
Routed MoE Down-Projections blk.*.ffn_down_exps.weight (All 48 Layers) 48 IQ4_NL (36) / Q4_0 (12) Non-linear codebook quantization for SSM layers, Q4_0 for anchor layers.
Routed MoE Gate/Up Projections blk.*.ffn_{gate,up}_exps.weight (All 48 Layers) 96 IQ3_XXS (72) / Q4_K (24) Calibrated with imatrix for maximum compactness, with Q4_K protection on anchor layers.
Residual Output Highway output_hc_{down,up}.weight 2 Q4_0 Low-rank residual highway projections at model termination.
Residual Output Highway Norm output_hc_norm.weight 1 F32 Final residual normalization anchor.

Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50 & RAM Streaming)

Hardware Target Offload Mode Generation Speed (Est.) Prompt Prefill Speed (Est.) Highlights
NVIDIA RTX 5080 / 5090 (Blackwell) Full GPU (-ngl 99) approx. 210 – 240 tok/s 2,600 – 3,500+ tok/s Blistering throughput on Hopper / Blackwell architectures
NVIDIA RTX 4090 (24GB GDDR6X) Full GPU (-ngl 99) 80 – 105+ tok/s 1,800 – 2,400+ tok/s Hybrid linear attention slashes prefill latency
NVIDIA RTX 3090 (24GB GDDR6) Full GPU (-ngl 99) 65 – 82+ tok/s 1,300 – 1,800+ tok/s Full 256k native window in VRAM
Workstation / Laptop (DDR4 / DDR5 RAM) Hybrid Offload (Few layers in VRAM) Hardware-dependent Hardware-dependent Zero AVX2 CPU stalls; efficient streaming from system RAM

The 24GB Miracle: Full 256K Context Runs In VRAM!

Qwen3.8-Flash-Coder-85GB APEX-I-MiniPlus-V2.1 fits deep context windows within standard 24GB and 32GB GPUs:

Context Length Model Weights KV Cache (q8_0) Compute Buffers Total GPU VRAM (Est.) Feasibility
32,768 (32k) 20.27 GiB 0.45 GiB 1.50 GiB 22.22 GiB Effortless fit on 24GB GPUs (RTX 3090 / 4090)
65,536 (64k) 20.27 GiB 0.85 GiB 1.65 GiB 22.77 GiB Full offload on 24GB GPUs
131,072 (128k) 20.27 GiB 1.45 GiB 1.90 GiB 23.62 GiB Full offload on 24GB GPUs
262,144 (256k) 20.27 GiB 2.60 GiB 2.40 GiB 25.27 GiB Full offload on 32GB (RTX 5090) or partial RAM offload

Recommended Configuration & Setup

llama-server \
  -m Qwen3.8-Flash-Coder-85GB.APEX-I-MiniPlus-V2.1.gguf \
  --jinja \
  -ngl 99 \
  --ctx-size 65536 \
  --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.00 \
  --port 8080

Recommended Generation Parameters

Hyperparameter Value Description / Creator Notice
Temperature 0.60 Recommended default for code generation & reasoning (--temp 0.6).
Top-P 0.95 Source sampling parameter (--top-p 0.95).
Top-K 20 Source sampling parameter (--top-k 20).
Min-P 0.00 Source sampling parameter (--min-p 0.00).
Template Engine --jinja Recommended official chat template flag.
Context Size 65536 High-throughput 64K context (up to native 256K).

CRITICAL ADVISORY FOR CODING WORKFLOWS: PREVENTING SYNTAX & TOKEN SWAPPING

In programming code, brackets ({, }), assignment operators (=), and indentation whitespace repeat constantly across multi-line structures.

Common Issue: Many local frontends (such as LM Studio defaults, Ollama, or web interfaces) ship with repeat_penalty set to 1.1 or 1.15. While this prevents loops in creative prose, applying repeat penalties to code artificially penalizes necessary syntax tokens. When the logit of { drops, the model is forced to emit the next closest mathematical token (= or [), resulting in character swapping or dropped/doubled whitespace.

Eliminating Character Swapping:

  1. Disable Repeat Penalties (Required for Code):
    • repeat_penalty: 1.0 (strictly disabled)
    • presence_penalty: 0.0
    • frequency_penalty: 0.0
  2. Calibrate Samplers:
    • temperature: 0.60 (or 0.20 - 0.30 for strict, deterministic code syntax)
    • min_p: 0.05 (prunes low-probability noise tokens effectively)
    • top_p: 0.95
    • top_k: 20
  3. Native Jinja Formatting: Always pass the --jinja flag so the tokenizer handles leading-space BPE tokens cleanly.

HARDENED AGENTIC CHAT TEMPLATE (JINJA)

The repository supports full Jinja chat templating with multi-level reasoning effort control:

  • low / minimal: Keeps internal thinking concise and focused strictly on immediate execution steps to minimize latency in automated loops.
  • medium (default): Balanced, structured reasoning process with standard analytical depth.
  • high / xhigh: Guides the model to formulate a clear implementation plan upfront before generating code.

7. Community Compute Fund & Priority Model Requests

Gold Ship dancing

All IsValorum quantizations will always remain completely free and open to the public without paywalls.

However, cloud GPU compute is expensive. If you find these builds valuable and would like to support the project or request a specific model architecture to be prioritized for the next MiniPlus/NanoPlus release, you can sponsor GPU compute time through Ko-fi:

Support on Ko-fi

(When supporting on Ko-fi, feel free to leave a note with your Hugging Face handle and the specific model you would like prioritized).

Downloads last month
1,804
GGUF
Model size
43B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF

Quantized
(6)
this model

Collections including IsValorum/Qwen3.8-Flash-Coder-APEX-I-MiniPlus-V2.1-GGUF