Instructions to use Qwen/Qwen3-32B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Qwen/Qwen3-32B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Qwen/Qwen3-32B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-32B") model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-32B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- HuggingChat
- Notebooks
- Google Colab
- Kaggle
- AMD Developer Cloud
- Local Apps Settings
- vLLM
How to use Qwen/Qwen3-32B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Qwen/Qwen3-32B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen3-32B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Qwen/Qwen3-32B
- SGLang
How to use Qwen/Qwen3-32B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Qwen/Qwen3-32B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen3-32B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Qwen/Qwen3-32B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen3-32B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Qwen/Qwen3-32B with Docker Model Runner:
docker model run hf.co/Qwen/Qwen3-32B
The correct way of fine-tuning on multi-turn trajectories
Looking at the qwen 3 chat template, the last assistant turn always has <think></think> tags, even in non-thinking mode, while the intermediate assistant turn never include reasoning traces and tags. This creates an asymmetry between the last assistant turn and all previous turns. And this asymmetry makes it unclear how to fine-tune this model on multi-turn trajectories: if one just does it by training on the whole trajectory with assistant turn masking, the intermediate turns will be OOD as they won't have thinking tags.
What's the recommended approach here? Should we just always train on the last turn or should we simply ignore this asymmetry?
Hi, we are facing the same problem while fine-tuning Qwen3-Base model. Is there a good reason of removing all the thinking blocks in the intermediate turns? Does thinking in intermediate turns harm the performance/training? Thanks!
You can fine-tune a multiturn trajectory by splitting it into multiple examples, and remove all
<think></think>tags of history turns.
So instead of train on the following,
train on the following, where blue ones are masked, red ones are trained.
Hope this helps!
it could be data leak, i.e. N-th trajectory appeared before N-1 traj -> so the model already knows the label

