Instructions to use inclusionAI/Ling-3.0-flash with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Inference
- HuggingChat
- Notebooks
- Google Colab
- Kaggle
How well it takes quantization...
#14
by MartinPatterson - opened
This comment has been hidden (marked as Resolved)
I was wondering this too so I ran some tests on my basic, no imatrix, GGUFs and from the looks of it I'd say quite well.
| File | Size | Strategy | PPL vs F16 |
|---|---|---|---|
Ling-3.0-flash-f16.gguf |
243,268 MiB (~237.6 GiB) | Full precision (baseline) | 4.3532 (reference) |
Ling-3.0-flash-SUPER-Q4_K_M.gguf |
82,476 MiB (~80.5 GiB) | Q8_0 signal path + Q4_K expert bulk | 4.3849 (+0.73%) |
Ling-3.0-flash-SUPER-Q3_K_M.gguf |
64,183 MiB (~62.7 GiB) | Q6_K signal path + Q3_K expert bulk | 4.5289 (+4.04%) |
Ling-3.0-flash-SUPER-Q2_K.gguf |
61,776 MiB (~60.3 GiB) | Q8_0 signal path + Q2_K expert bulk | 4.9592 (+13.92%) |