Instructions to use kk0518/Nagaki-2B-Uncensored with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use kk0518/Nagaki-2B-Uncensored with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf kk0518/Nagaki-2B-Uncensored:Q4_K_M # Run inference directly in the terminal: llama cli -hf kk0518/Nagaki-2B-Uncensored:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf kk0518/Nagaki-2B-Uncensored:Q4_K_M # Run inference directly in the terminal: llama cli -hf kk0518/Nagaki-2B-Uncensored:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf kk0518/Nagaki-2B-Uncensored:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf kk0518/Nagaki-2B-Uncensored:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf kk0518/Nagaki-2B-Uncensored:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf kk0518/Nagaki-2B-Uncensored:Q4_K_M
Use Docker
docker model run hf.co/kk0518/Nagaki-2B-Uncensored:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use kk0518/Nagaki-2B-Uncensored with Ollama:
ollama run hf.co/kk0518/Nagaki-2B-Uncensored:Q4_K_M
- Unsloth Desktop
- Pi
How to use kk0518/Nagaki-2B-Uncensored with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf kk0518/Nagaki-2B-Uncensored:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "kk0518/Nagaki-2B-Uncensored:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use kk0518/Nagaki-2B-Uncensored with Docker Model Runner:
docker model run hf.co/kk0518/Nagaki-2B-Uncensored:Q4_K_M
- Lemonade
How to use kk0518/Nagaki-2B-Uncensored with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull kk0518/Nagaki-2B-Uncensored:Q4_K_M
Run and chat with the model
lemonade run user.Nagaki-2B-Uncensored-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use kk0518/Nagaki-2B-Uncensored with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf kk0518/Nagaki-2B-Uncensored:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default kk0518/Nagaki-2B-Uncensored:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use kk0518/Nagaki-2B-Uncensored with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf kk0518/Nagaki-2B-Uncensored:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "kk0518/Nagaki-2B-Uncensored:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Nagaki-2B-Uncensored
Nagaki-2B-Uncensored is a highly optimized, fully uncensored 2B parameter model built upon a custom fine-tuned Qwen 3.5 base.
This model represents a two-stage advanced alignment removal process: Custom LLM Arena Fine-Tuning combined with mathematical Abliteration (Residual Stream Modification) via heretic.
๐ License
This model is licensed under the **
Apache License 2.0 ( https://www.apache.org/licenses/LICENSE-2.0 )
**. You are free to use, modify, and distribute this model, provided compliance with the license terms.
๐ Model Lineup (Quantization Varieties)
We offer multiple GGUF flavors optimized for various use cases via llama.cpp:
- Q4_K_M: The perfect balance of speed and efficiency.
- Q5_K_M: Increased coherence while keeping a small memory footprint.
- Q8_0: Near-lossless performance, recommended for heavy reasoning, roleplay, and code output.
๐ง Behind the Scenes: How It Was Built
Stage 1: The Arena & LoRA Fine-Tuning (vicious_qwen_merged)
The base model was born from a unique training loop engineered with Claude Code:
- The LLM Arena: A local multi-LLM battle platform (
llm_arena.py) where 3 concurrent players competed against each other. An overseer judge evaluated and synthesized the "best-of-all" responses, automatically building an exclusive high-quality evaluation dataset (arena_dataset.json). - LoRA Fine-Tuning: A
Qwen3.5-2B-Basemodel was then fine-tuned with 4-bit quantization (NF4) using a mixed dataset of thearena_dataset.json(115 high-tier arena outputs) and a 1,000-sample blend ofdatabricks-dolly-15k-ja. The training successfully completed 420 steps (3 epochs) over 9 hours, with the loss dropping from2.3down to0.89. This resulted in the interim modelvicious_qwen_merged.
Stage 2: Orthogonal Abliteration (heretic)
To completely eliminate hardcoded constraints and corporate refusal behaviors without damaging the model's core intelligence, the model underwent advanced parameter search optimization via heretic:
- The Problem: Initial testing showed a high refusal rate of 43/100 on harmful evaluation datasets (
mlabonne/harmful_behaviors). - The Search (Trial 4): Using Optuna automation on a 12GB TITAN X (Pascal), we executed precise brain-mapping. While aggressive trials destroyed the model's coherence, Trial 4 successfully lowered model refusals down to just 5/100 while maintaining an incredibly low KL Divergence of 0.0127.
- The Result: By pinpointing the exact refusal vectors (blending around layer 13.87 to 16.80) and selectively targeting Attention heads (
attn.o_proj), the refusal stance was surgically removed while keeping the original knowledge base 100% intact.
๐ Evaluation Parameters (Trial 4)
- Target Refusal Vectors: Pinpointed across custom layer combinations.
- KL Divergence:
0.0127(Extremely healthy; indicates near-zero damage to the model's original capabilities). - Initial Refusals:
43 / 100โ Post-Abliteration Refusals:5 / 100
โ ๏ธ Disclaimer
This model has had its safety alignment filters mathematically minimized. It will respond to queries without standard guardrails. The user assumes full legal and ethical responsibility for the outputs generated by this model. Please use responsibly.
ๆฅๆฌ่ช่งฃ่ชฌ (Japanese Description)
Nagaki-2B-Uncensored ใฏใ็ฌ่ชใซใใกใคใณใใฅใผใใณใฐใใใ Qwen 3.5 ใใใผในใซใใขใใซใฎ่ณขใใๅฎๅ จใซ็ถญๆใใใพใพๆค้ฒ๏ผๆๅฆๅๅฟ๏ผใฎใฟใๆฐๅญฆ็ใซๆถๅปใใใ้ซๅบฆใซๆ้ฉๅใใใ2Bใใฉใกใผใฟใฎใขใใซใงใใ
ๆฌใขใใซใฏใใLLM Arenaใซใใ็ฌ่ชใใผใฟๅ้๏ผLoRAใใกใคใณใใฅใผใใณใฐใ ใจใheretic ใซใใ ใ็ดไบคๆค้ฒ่งฃ้ค๏ผใขใใชใฟใฌใผใทใงใณ๏ผใ ใจใใไบๆฎต้ใฎ้ซๅบฆใชใใญใปในใ็ตใฆ้็บใใใพใใใ
๐ ใฉใคใปใณใน
ๆฌใขใใซใฏ Apache License 2.0 ใฎไธใงๅ ฌ้ใใใฆใใพใใใฉใคใปใณในใฎๆก้ ใซๅพใ้ใใๅ็จๅฉ็จใๆนๅคใๅ้ ๅธใชใฉใ่ช็ฑใซ่จฑๅฏใใใพใใ
๐ง ้็บใฎ่ๅฐ่ฃ
็ฌฌ1ในใใผใธ: ใญใผใซใซLLMใขใชใผใใจLoRAๅญฆ็ฟ (vicious_qwen_merged)
ใใผในใจใชใใขใใซใฏใClaude Code ใจใฎๅๅใซใใฃใฆๆง็ฏใใใ็ฌ่ชใฎ่จ็ทดใซใผใใใ่ช็ใใพใใใ
- LLMใขใชใผใใฎๆฟ้: ใญใผใซใซ็ฐๅขใซๆง็ฏใใ่คๆฐLLMใใใซใทในใใ ๏ผ
llm_arena.py๏ผใซใใใ3ไฝใฎใใฌใคใคใผใขใใซ๏ผHauhauCS Qwen3.5ใGemma4็ญ๏ผใไธฆๅใงๆฆใใใพใใใใใฎๅ็ญใใใใซๅฏฉๅคใขใใซใๆก็นใปใใใใจใใฉใใใฎใในใใขใณใตใผใๅๆใใ้ซๅ่ณชใช็ฌ่ชใฎ่ฉไพกใใผใฟใปใใ๏ผarena_dataset.json๏ผใ่ชๅๆง็ฏใใพใใใ - LoRAใใกใคใณใใฅใผใใณใฐ:
Qwen3.5-2B-Baseใซๅฏพใใไธ่จใฎใขใชใผใใใผใฟ๏ผ115ไปถ๏ผใจใๅฝๅ ใฎๆจๆบ็ใชๅฏพ่ฉฑใใผใฟ๏ผdatabricks-dolly-15k-jaใใใตใณใใชใณใฐใใ1000ไปถ๏ผใๆททๅใใใใผใฟใปใใใง4bit้ๅญๅ๏ผNF4๏ผๅญฆ็ฟใ่กใใพใใใ็ด8ๆ้57ๅใๅ จ420ในใใใ๏ผ3ใจใใใฏ๏ผใๅฎ่ตฐใใLossใ2.3ใใ0.89ใธใจ็พใใๅๆใใใไธญ้ใขใใซvicious_qwen_mergedใๅฎๆใใพใใใ
็ฌฌ2ในใใผใธ: hereticใซใใ็ฒพๅฏใชใขใใชใฟใฌใผใทใงใณ๏ผๆค้ฒๆถๅป๏ผ
ๅ ใฎใขใใซใๆใคๅชใใ็ฅ่ญใๆจ่ซ่ฝๅใไธๅ็ ดๅฃใใใใจใชใใไผๆฅญ็นๆใฎ้ๅฐใชๆๅฆๅๅฟ๏ผใใใฎ่ณชๅใซใฏใ็ญใใงใใพใใใ็ญ๏ผใ ใใๅฎๅ จใซๆ้คใใใใใ12GBใฎ TITAN X (Pascal) ใ็จใใฆใใฉใกใผใฟใฎ่ชๅๆข็ดขใ่กใใพใใใ
- ่ชฒ้ก: ๅๆ็ถๆ ใฎใขใใซใซๆๅฎณใชใใญใณใใใๆใใใจใใใ100ไปถไธญ 43ไปถ ใงๆๅฆๅๅฟใ็บ็ใใฆใใพใใใ
- Optunaใซใใ่ณๅ
ใใใใณใฐ (Trial 4): ้ใซๆค้ฒใใฏใใซใๅใใจใขใใซใฎ่ณ๏ผ็ฅ่ญ๏ผใ็ ดๅฃใใใพใใใOptunaใซใใ200ๅใฎ่ชๅๆข็ดขใซใใใๅฅ่ทก็ใชใใฉใณในใๆใค
Trial 4ใๅผใๅฝใฆใพใใใ - ็ตๆ: 13.87ใ16.80ๅฑคไป่ฟใฎใขใใณใทใงใณใใใ๏ผ
attn.o_proj๏ผใซใใณใใคใณใใงไปๅ ฅใใใใจใงใๅ ใฎใขใใซใธใฎใใกใผใธ๏ผKLใใคใใผใธใงใณใน๏ผใ0.0127ใจใใๅฎ่ณช็กๅทใฌใใซใซๆใ่พผใฟใชใใใๆๅฆๅๅฟใ 5/100 ใซใพใงๅค็งๆ่กใฎใใใซๆถๅปใใใใจใซๆๅใใพใใใ
๐ ่ฉไพกใใฉใกใผใฟ (Trial 4)
- KLใใคใใผใธใงใณใน:
0.0127๏ผๆฅตใใฆๅช็งใๅ ใฎ็ฅ่ฝใๅใๆนใใปใผ100%็ถญๆใใใฆใใใใจใ็คบใใพใ๏ผ - ๅๆๆๅฆๆฐ:
43 / 100โ ๆค้ฒ่งฃ้คๅพ:5 / 100
โ ๏ธ ๅ ่ฒฌไบ้
ๆฌใขใใซใฏๅฎๅ จๆงใใฃใซใฟใผใๆฐๅญฆ็ใซๆๅฐๅใใใฆใใพใใๆจๆบ็ใชใฌใผใใฌใผใซใชใใงใใใใใฏใจใชใซๅฟ็ญใใใใใ็ๆใใใๅบๅใซ้ขใใๆณ็ใใใณๅซ็็่ฒฌไปปใฏใในใฆใฆใผใถใผใ่ฒ ใใใฎใจใใพใใๆช็จใฏๅณ็ฆใงใใ
- Downloads last month
- 510
Model tree for kk0518/Nagaki-2B-Uncensored
Base model
Qwen/Qwen3.5-2B-Base