How well it takes quantization...

#14
by MartinPatterson - opened
This comment has been hidden (marked as Resolved)

I was wondering this too so I ran some tests on my basic, no imatrix, GGUFs and from the looks of it I'd say quite well.

File Size Strategy PPL vs F16
Ling-3.0-flash-f16.gguf 243,268 MiB (~237.6 GiB) Full precision (baseline) 4.3532 (reference)
Ling-3.0-flash-SUPER-Q4_K_M.gguf 82,476 MiB (~80.5 GiB) Q8_0 signal path + Q4_K expert bulk 4.3849 (+0.73%)
Ling-3.0-flash-SUPER-Q3_K_M.gguf 64,183 MiB (~62.7 GiB) Q6_K signal path + Q3_K expert bulk 4.5289 (+4.04%)
Ling-3.0-flash-SUPER-Q2_K.gguf 61,776 MiB (~60.3 GiB) Q8_0 signal path + Q2_K expert bulk 4.9592 (+13.92%)

Sign up or log in to comment