Instructions to use VikramPal/Qwen3.5-35B-A3B-IntentRouter with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use VikramPal/Qwen3.5-35B-A3B-IntentRouter with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="VikramPal/Qwen3.5-35B-A3B-IntentRouter") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("VikramPal/Qwen3.5-35B-A3B-IntentRouter") model = AutoModelForCausalLM.from_pretrained("VikramPal/Qwen3.5-35B-A3B-IntentRouter", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use VikramPal/Qwen3.5-35B-A3B-IntentRouter with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "VikramPal/Qwen3.5-35B-A3B-IntentRouter" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VikramPal/Qwen3.5-35B-A3B-IntentRouter", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/VikramPal/Qwen3.5-35B-A3B-IntentRouter
- SGLang
How to use VikramPal/Qwen3.5-35B-A3B-IntentRouter with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "VikramPal/Qwen3.5-35B-A3B-IntentRouter" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VikramPal/Qwen3.5-35B-A3B-IntentRouter", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "VikramPal/Qwen3.5-35B-A3B-IntentRouter" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VikramPal/Qwen3.5-35B-A3B-IntentRouter", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use VikramPal/Qwen3.5-35B-A3B-IntentRouter with Docker Model Runner:
docker model run hf.co/VikramPal/Qwen3.5-35B-A3B-IntentRouter
Qwen3.5-35B-A3B Intent Router
A general intent router: it reads a user turn plus an intent catalog supplied in the prompt and returns exactly one intent id from that catalog. The catalog is an input, not a trained-in label set, so the same checkpoint routes for a catalog it has never seen -- which is what the held-out BANKING77 column below measures.
Fine-tuned from Qwen/Qwen3.5-35B-A3B.
Read this before you load it
Three things about this checkpoint will produce wrong results silently if you do not know them.
1. It is text-only. The base is a multimodal qwen3_5_moe checkpoint. What is published
here is the text tower alone -- model_type is qwen3_5_moe_text, and the vision config
and the MTP speculative-decoding head are not present. 34,660,610,688 parameters against the base
checkpoint's 35,951,822,704. If you need vision or MTP, use the base model.
2. Load it as a normal transformers model.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained(
"VikramPal/Qwen3.5-35B-A3B-IntentRouter", dtype=torch.bfloat16, device_map="auto",
trust_remote_code=True,
experts_implementation="eager", # grouped_mm needs sm_90; drop this on Hopper+
)
tok = AutoTokenizer.from_pretrained("VikramPal/Qwen3.5-35B-A3B-IntentRouter")
3. vLLM cannot serve this checkpoint. vLLM refuses fused MoE weights at layer 0. Use
transformers.
How to prompt it
The system prompt carries the catalog; the user turn carries the utterance. The model was
trained on a fixed system template (published in this repo as system_template.txt) and
renders through the chat template with enable_thinking=False. Scoring at eval time used
add_special_tokens=False because the template already emits them -- passing both prepends a
second BOS and measures a different model.
messages = [
{"role": "system", "content": SYSTEM_TEMPLATE.format(catalog="\n".join(intent_ids))},
{"role": "user", "content": "my card hasn't turned up yet"},
]
text = tok.apply_chat_template(messages, tokenize=False,
add_generation_prompt=True, enable_thinking=False)
enc = tok(text, return_tensors="pt", add_special_tokens=False).to(model.device)
out = model.generate(**enc, max_new_tokens=24, do_sample=False)
print(tok.decode(out[0, enc.input_ids.shape[1]:], skip_special_tokens=True))
# -> card_arrival
Greedy decoding, and 24 new tokens is enough for every label in these catalogs. The router
answers with a bare intent id, or out_of_scope when the catalog does not cover the turn,
or a clarifying question when the turn conjoins two intents.
Results
1,500 held-out items, 150 per (split, catalog) group, greedy decoding, strict exact match against the gold intent id.
| arm | strict | right id present | invented ids | truncated at 24 tokens |
|---|---|---|---|---|
| Qwen3.5-35B-A3B (no fine-tune) | 62.5% | 62.5% | 9.3% | 14/1500 |
| this router, bf16 | 91.1% | 91.1% | 0.3% | 0/1500 |
| this router, DynQuant 4-bit | 90.6% | 90.6% | 0.3% | 0/1500 |
| this router, DynQuant 3-bit | 79.1% | 80.6% | 3.6% | 41/1500 |
Invented ids are answers that are not in the catalog the model was given -- the failure mode that makes a router unusable downstream, because the caller has no branch for them.
Strict requires the reply to be exactly the intent id and nothing else; right id present also accepts the id surrounded by other text. The two are identical for Qwen3.5-35B-A3B (no fine-tune), this router, bf16, this router, DynQuant 4-bit, so those strict numbers are accuracy outright. They separate on this router, DynQuant 3-bit (22 of 1500 items), which reached the correct id and then kept writing -- a note, a caveat, or a markdown table. Whether that counts as a loss is the caller's choice: a router that reads the first line of the reply recovers those items, one that requires a bare label does not. Both numbers are given so neither reading has to be taken on trust.
Paired McNemar against the un-fine-tuned base model on the same 1,500 items (18 items where only that arm was right, 447 where only this one was): +28.60 points, exact binomial p = 2.53e-108. The test is paired because both arms answer the identical items; the ~1035 items they agree on carry no information about which is better.
By catalog
| catalog | intents | in training | strict |
|---|---|---|---|
banking77 |
77 | held out | 84.3% |
clinc150 |
151 | yes | 98.3% |
hwu68 |
67 | yes | 89.3% |
massive60 |
60 | yes | 87.3% |
mtop117 |
113 | yes | 96.0% |
banking77 is never seen in training -- neither its utterances nor its 77 intent ids. It is the column that says whether this is a router or a 77-way classifier wearing one.
By item kind
| kind | what it tests | strict |
|---|---|---|
clarify_conj |
two intents conjoined -- should ask, not guess | 97.3% |
cs_oos |
genuinely out of scope for any catalog intent | 92.0% |
cs_removed |
the correct intent was deleted from the catalog | 86.4% |
multiturn |
the intent is only resolvable from prior turns | 90.4% |
normal |
a plain utterance with its intent in the catalog | 90.8% |
same_intent_conj |
two clauses, one intent | 91.6% |
By language
| de | en | es | fr | hi | th |
|---|---|---|---|---|---|
| 87.4% | 91.1% | 94.3% | 92.9% | 90.6% | 90.5% |
Training
LoRA (r=32, alpha=64, dropout=0.05) on bf16 base weights -- not QLoRA; the base is not
quantized during training, so the harvested gradient signal describes the tensors that are
actually quantized afterwards. Adapters on every attention and MLP projection including the
linear-attention in_proj_* family, merged into the base before quantization.
| train | 164,473 examples |
| validation | 26,693 examples |
| rule-behaviour set | 4,556 examples (clarify / out-of-scope / multi-turn / catalog-removal) |
| held-out test | 46,605 examples, of which 1,500 scored |
| languages | en, de, es, fr, hi, th |
| catalogs in training | clinc150, hwu68, massive60, mtop117 |
Training data is assembled from five public intent datasets, each turned into catalog-in-prompt form and augmented with the rule behaviours the router has to get right (refusing to guess when the intent was removed from the catalog, asking when two intents are conjoined, staying silent about intents that are not offered).
| dataset | intents | used for | licence |
|---|---|---|---|
| MASSIVE | 60 | train+test | CC BY 4.0 |
| CLINC150 | 151 | train+test | CC BY 3.0 |
| BANKING77 | 77 | test only (held out) | CC BY 4.0 |
| HWU64 | 67 | train+test | CC BY 4.0 |
| MTOP | 113 | train+test | CC BY-SA 4.0 |
BANKING77 appears only in the test split -- it is the generalization measurement, and training on it would destroy the only evidence that the catalog is really an input.
Limitations
- Text only. No vision tower, no MTP head (see the top of this card).
- Not servable by vLLM. vLLM's fused-MoE guard rejects this architecture at layer 0.
- Six languages. en/de/es/fr/hi/th. Other languages are untested and the catalog ids are English regardless of the utterance language.
- Catalog size. Tested at 60-151 intents. Much larger catalogs will not fit the context the same way and are unmeasured here.
- One intent per turn. Conjoined intents are trained to produce a clarifying question, not two labels.
experts_implementation="eager"is needed below sm_90. The default grouped-MoE path requires Hopper or newer.
Citation
The quantization method:
@misc{dynquant2026,
title = {DynQuant: Dynamic-Signal Quantization for Extreme LLM Compression},
author = {Kamboj, Vikram Pal and Kour, Manpreet},
year = {2026}
}
The base model is Qwen/Qwen3.5-35B-A3B; please cite Qwen as well, and the five source datasets listed above.
- Downloads last month
- 520