Q4 vs Q5 vs Q6 vs Q8
A 32K planning comparison for Qwen3.8-27B on 24GB. Memory estimates are useful; quality claims remain unknown until the same evaluation is run across every artifact.
| Quantization | Weight estimate | Peak range | 24GB verdict | Quality evidence |
|---|---|---|---|---|
| Q4_K_M | 16.7 GB | 21.4 GB–22.3 GB | fits | Not yet tested |
| Q5_K_M | 19.8 GB | 24.7 GB–25.8 GB | does not fit | Not yet tested |
| Q6_K | 22.9 GB | 28.1 GB–29.3 GB | does not fit | Not yet tested |
| Q8_0 | 29.5 GB | 35.2 GB–36.8 GB | does not fit | Not yet tested |
| BF16 | 55.6 GB | 63.2 GB–66.2 GB | does not fit | Not yet tested |
INTERPRETATION
Fit and quality are different questions
Lower-bit weights create more memory headroom, but that does not establish an acceptable quality trade-off. The lab will only publish a comparative quality statement after the evaluation prompts, scoring method, runtime and generation settings are held constant.