Running 4.08k The Ultra-Scale Playbook 🌌 4.08k The ultimate guide to training LLM on large GPU Clusters
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads Paper • 2401.10774 • Published Jan 19, 2024 • 60
mistralai/Mistral-7B-Instruct-v0.1 Text Generation • 7B • Updated Jul 24, 2025 • 145k • • 1.85k