minimax_h3_ref2va_pruned_w4a8_mixed.safetensors vs minimax_h3_ref2va_pruned_w6a8_g32.safetensors
#50
by makisekurisu-jp - opened
Are these two models only different in weight quantization
— one using 4-bit quantization and the other using 6-bit quantization
— while both use INT8 activations, meaning that their actual accuracy is essentially the same?