What is established
- Parameters
- 27.78B
- Architecture
- Qwen3_5ForConditionalGeneration
- Transformer blocks
- 64
- Full-attention blocks
- 16
- Native context
- 262,144 tokens
- FP16 KV cache
- 64 KiB / token
- Repository revision
1d4bf0f2ff60…- License
- Apache-2.0
Compare local memory requirements, choose a setup and find the exact tested GGUF. Calculated fit is separate from the observed RTX 4090 results.
These scenarios use Ollama, 4-bit model files, 8,192 tokens of context and one chat session. NVIDIA estimates assume 64 GiB of system RAM and no attached display. Mac estimates reserve at least 6 GiB for macOS. Open a setup to change those assumptions or compare runtimes.
| Your hardware | Model files | Estimated peak / available | Fit / runtime review |
|---|---|---|---|
| 16GB NVIDIA GPUCheck this setup → | Q4_K_MOllama · 15.7 GiB weights | 17.5 GiB15.5 GiB available | does not fitRuntime documented |
| 24GB NVIDIA GPUCheck this setup → | Q4_K_MOllama · 15.7 GiB weights | 17.5 GiB23.5 GiB available | comfortableRuntime documented |
| 32GB NVIDIA GPUCheck this setup → | Q4_K_MOllama · 15.7 GiB weights | 17.5 GiB31.5 GiB available | comfortableRuntime documented |
| 32GB Apple SiliconCheck this setup → | Q4_K_MOllama · 15.7 GiB weights | 18.8 GiB26.0 GiB available | comfortableRuntime documented |
| 64GB Apple SiliconCheck this setup → | Q4_K_MOllama · 15.7 GiB weights | 18.8 GiB56.3 GiB available | comfortableRuntime documented |
Memory estimates are not speed tests or a guarantee of fit. System RAM is not added to dedicated VRAM; partial CPU offload is outside these scenarios. Q5, Q6 and Q8 are not verified for this model in the planner.
Mac setups start with Ollama. MLX LM remains available for comparison, but its reviewed Qwen3.8 implementation does not enforce the requested cache-size limit, so those commands are withheld.
Source checks are dated below. Calculator inputs and recorded benchmarks have separate review dates.
Last check: 2026-09-27 03:43 UTC · 3 changes for review · 0 unavailable
Inspect source health →1d4bf0f2ff60…15.9 GiB on disk · unsloth/Qwen3.8-27B-GGUF
This is the file used in the linked llama.cpp RTX 4090 tests. The Ollama and MLX planner artifacts have their own file sizes and identifiers; the test does not measure their speed.
Inspect the pinned GGUF file ↗Revision: f1bfb127c64f7072bdd2cad55f258b9c8b2910fe
SHA-256: 7e78da5d7e3ae28d178121f58646953305f3e5bd3cb46f4a75584e8b6c6fe169
Repeat the fit boundary across Q5, Q6 and Q8 using pinned artifacts.
Other GPUs, Macs and CPU-only systems need separate receipts.
A pre-registered evaluation across identical prompts and settings.
Evidence first. Estimates labelled. Corrections preserved.