Instructions to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF # Run inference directly in the terminal: llama cli -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF # Run inference directly in the terminal: llama cli -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF # Run inference directly in the terminal: ./llama-cli -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Use Docker
docker model run hf.co/IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
- LM Studio
- Jan
- Ollama
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF with Ollama:
ollama run hf.co/IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
- Unsloth Desktop
- Pi
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF with Docker Model Runner:
docker model run hf.co/IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
- Lemonade
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Run and chat with the model
lemonade run user.Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Navigation Index
Independently computed on private cloud clusters. If this handcrafted release saves you VRAM and runs faster on your GPU, consider fueling the community compute fund on Ko-fi.
- Optimization History & Transparency Notice
- Quality Spectrum: APEX-I-NanoPlus vs. Standard Flat Quantizations
- Model Files & Technical Specifications
- Surgical Tensor Quantization Map (Audited from GGUF)
- Inference Quickstart
- CRITICAL: Coding Syntax & Repeat Penalty Advisory (Preventing Character Swapping)
- Hardened Agentic Chat Template & Reasoning Effort
- Optional Support
EXPERIMENTAL PRE-RELEASE NOTICE: ENGLISH-ONLY CODING SPECIALIST
This model suite is quantized from Jab1718/qwen3.8-flash-coder-85gb-bf16, which is an intermediate experimental slice created using
moe-slice(352 out of 512 routed experts were permanently pruned exclusively against English Python and SWE-bench calibration datasets).
- English Coding Only: This model is strictly designed for programming, code completion, refactoring, and agentic tool-calling in English.
- Severe Multilingual & General Degradation: Because conversational and multilingual experts were pruned and the upstream author has not yet released the recovery fine-tuning pass, this model severely degrades and outputs broken text in languages other than English (e.g., Spanish, French, German, etc.) or in general chit-chat.
- Incompatible with Strata Engine: This model uses a 160-expert layout with decoupled n-gram tables; it is not compatible with Strata Engine (which requires the 512-expert monolith and 51B PLE tables). Run using stock
llama.cpp(llama-server) or LM Studio.
Qwen3.8-Flash-Coder APEX-I-NanoPlus GGUF
The Next-Generation Frontier MoE · Extreme 18GB Footprint · Fast System RAM Streaming & Massive Context on 16GB–24GB VRAM
EXPLORE THE COMPLETE QWEN3.8 FLASH CODER LINEUP
These are complementary APEX-I releases, not alternate downloads of the same model:
- Qwen3.8-Flash-Coder APEX-I-MiniPlus-V2.1 — specialist agentic coding MoE tuned for Q5–Q6 quality (21.77 GB / 3.45 BPW).
- Qwen3.8-Flash-Coder APEX-I-NanoPlus — ultra-compact footprint achieving solid Q4 quality (18.34 GB / 2.90 BPW).
- Qwen3.8-Flash-Coder-85GB Lossless BF16 GGUF — uncompressed reference baseline (85.30 GB / 16.00 BPW).
EMPIRICAL BENCHMARK & QUALITY COMPARISON
Quantization Specification File Size (Disk) Memory Footprint (RAM/VRAM) Average BPW WikiText-2 Perplexity Delta PPL vs BF16 (%) Quality Tier Equivalent Uncompressed BF16 Reference 85.30 GB(79.44 GiB)79.44 GiB16.00 BPW 30.0975 +/- 0.1200Baseline (0.00%) Lossless Reference Baseline APEX-I-MiniPlus V2.1 21.77 GB(20.27 GiB)20.27 GiB3.45 BPW 30.1495 +/- 1.0089+0.0520 (+0.17%)Q5_K_L / Q6_K Tier Boundary APEX-I-NanoPlus (CURRENT) 18.34 GB(17.08 GiB)17.08 GiB2.90 BPW 34.4199 +/- 1.1591+4.3224 (+14.36%)Solid Q4_K_M Tier Standard Flat Q3_K_S 20.41 GB 19.01 GiB 3.10 BPW aprox. 30.75 - 31.20 +0.65 a +1.10 (+2.9%) High syntax degradation Generic APEX Mini (IQ2_S) 17.73 GB 16.51 GiB 2.50 BPW aprox. 31.60 - 33.10+ +1.50 a +3.00+ (+7.5%) Severe reasoning breakdown
Model Files & Technical Specifications
| File Name | File Size | Memory Footprint | BPW | Description |
|---|---|---|---|---|
Qwen3.8-Flash-Coder-85GB.APEX-I-NanoPlus.gguf |
18.34 GB (17.08 GiB) |
17.08 GiB |
2.90 BPW | Ultra-compact agentic coding MoE achieving solid Q4 quality with massive context capability |
- Base Model: Jab1718/qwen3.8-flash-coder-85gb-bf16
- Parameters: 45.8B total (approx. 3.7B active per token: 10 routed experts + 1 shared expert + dense backbone; 4.9B with vocabulary embeddings)
- Architecture:
Qwen4ExpForCausalLM(48 hybrid layers, 160 MoE experts with 10 active) - Context Length: 262,144 tokens (native 256K)
Surgical Tensor Quantization Map (Audited from GGUF)
| Layer Group | Sub-Component / Tensor | Qty | Precision | Engineering Rationale |
|---|---|---|---|---|
| Global Output Head | output.weight |
1 | Q6_K |
Armored in high-precision Q6_K to preserve token classification. |
| Global Embeddings | token_embd.weight |
1 | Q3_K |
Compact embedding representation across 248k vocabulary. |
| All Normalizations | output_norm, attn_norm, ffn_norm, hc_norm, ssm_norm |
146 | F32 |
100% uncompressed numerical stability across all 48 layers. |
| Expert Routers | blk.*.ffn_gate_inp, ffn_gate_inp_shexp |
96 | F32 |
100% uncompressed routing fidelity across 160 experts. |
| Attention Gates | blk.*.attn_gate.weight (36 SSM Layers) |
36 | Q8_0 |
High-precision attention gating across hybrid DeltaNet recurrence layers. |
| Linear Attention Projections | blk.*.attn_qkv.weight (36 SSM Layers) |
36 | Q3_K |
Efficient 3-bit quantization for SSM attention state inputs. |
| SSM Linear Output | blk.*.ssm_out.weight (36 SSM Layers) |
36 | Q5_K |
5-bit precision for linear state-space recurrence output. |
| Recurrent SSM Parameters | blk.*.ssm_{a,alpha,beta,conv1d,dt,norm} |
216 | F32 |
Guarded in uncompressed FP32 to prevent DeltaNet recurrent state drift. |
| Periodic Sparse Attention | blk.{3,7,...}.attn_{q,k,v}.weight (12 Layers) |
36 | Q4_K |
Quadratic attention checkpoints for deep retrieval. |
| Periodic Sparse Attention Output | blk.{3,7,...}.attn_output.weight (12 Layers) |
12 | Q6_K |
Armored attention output projection over deep context. |
| QSA Sparse Attention Indexer | blk.{3,7,...}.indexer.{q,k}_proj.weight |
24 | BF16 |
High-fidelity sparse indexer projections for query-stream attention routing. |
| QSA Indexer Norms | blk.{3,7,...}.indexer.{q,k}_norm.weight |
24 | F32 |
Uncompressed indexer layer normalizations. |
| Shared Foundation Experts | blk.*.ffn_down_shexp.weight (All 48 Layers) |
48 | Q5_0 |
ne0=640 in standard block-32; zero divisibility crashes. |
| Shared Foundation Experts | blk.*.ffn_{gate,up}_shexp.weight (All 48 Layers) |
96 | Q5_K |
Preserves core coding knowledge active on 100% of tokens. |
| Highway Connections (Down/Inject) | blk.*.hc_{attn,ffn}_{down,inject}.weight |
192 | Q6_K |
High-fidelity residual highway bypass. |
| Highway Connections (Up) | blk.*.hc_{attn,ffn}_up.weight |
96 | Q5_0 |
ne0=320 in standard block-32 format. |
| Routed MoE Down-Projections | blk.*.ffn_down_exps.weight (All 48 Layers) |
48 | IQ4_NL (36) / Q4_0 (12) |
Non-linear codebook quantization for SSM layers, Q4_0 for anchor layers. |
| Routed MoE Gate/Up Projections | blk.*.ffn_{gate,up}_exps.weight (All 48 Layers) |
96 | IQ2_S (36) / IQ2_XXS (36) / IQ3_XXS (24) |
Scaled expert density: sub-2.5 BPW on deep layers, IQ3_XXS on anchor layers. |
| Residual Output Highway | output_hc_{down,up}.weight |
2 | Q4_0 |
Low-rank residual highway projections at model termination. |
| Residual Output Highway Norm | output_hc_norm.weight |
1 | F32 |
Final residual normalization anchor. |
Inference Quickstart
llama-server \
-m Qwen3.8-Flash-Coder-85GB.APEX-I-NanoPlus.gguf \
--jinja \
-ngl 99 \
--ctx-size 65536 \
--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.00 \
--port 8080
CRITICAL ADVISORY FOR CODING WORKFLOWS: PREVENTING SYNTAX & TOKEN SWAPPING
Disable repeat penalties (
repeat_penalty: 1.0,presence_penalty: 0.0,frequency_penalty: 0.0) and use--jinjato avoid syntax bracket substitutions.
7. Community Compute Fund & Priority Model Requests
All IsValorum quantizations will always remain completely free and open to the public without paywalls.
However, cloud GPU compute is expensive. If you find these builds valuable and would like to support the project or request a specific model architecture to be prioritized for the next MiniPlus/NanoPlus release, you can sponsor GPU compute time through Ko-fi:
(When supporting on Ko-fi, feel free to leave a note with your Hugging Face handle and the specific model you would like prioritized).
- Downloads last month
- 1,431
We're not able to determine the quantization variants.
Model tree for IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Base model
Qwen/Qwen3.8-Flash-Next