Text Generation
Transformers
PyTorch
Safetensors
bloom
Eval Results (legacy)
text-generation-inference
Instructions to use bigscience/bloom-3b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bigscience/bloom-3b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="bigscience/bloom-3b")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("bigscience/bloom-3b") model = AutoModelForCausalLM.from_pretrained("bigscience/bloom-3b") - Notebooks
- Google Colab
- Kaggle
- Local Apps
- vLLM
How to use bigscience/bloom-3b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bigscience/bloom-3b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bigscience/bloom-3b", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/bigscience/bloom-3b
- SGLang
How to use bigscience/bloom-3b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "bigscience/bloom-3b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bigscience/bloom-3b", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "bigscience/bloom-3b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bigscience/bloom-3b", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use bigscience/bloom-3b with Docker Model Runner:
docker model run hf.co/bigscience/bloom-3b
Commit History
Add evaluation (#8) 68331cd
341b tokens model (#7) 538fcc1
Update README.md 7fac334
Younes Belkada commited on
Add correct pipeline tag (#2) 51a38df
ybelkada commited on
Update README.md (#1) df69d83
ybelkada commited on
Copy+Paste from updated model card, + warning. b6b0418
Update README.md d8263c9
Update date on Model Card 80fbc65
Younes Belkada commited on
Update README.md 030b11e
Added intermediary checkpoint warning 4304d27
Teven Le Scao commited on
Update Model Card 5118142
Younes Belkada commited on
initial commit 58f9b04
Younes Belkada commited on