Kingy AI
Explore
Work with Kingy
Menu
MODEL RECORD · REQUIREMENTS + TESTED FILE

Qwen3.8-27B requirements

Compare local memory requirements, choose a setup and find the exact tested GGUF. Calculated fit is separate from the observed RTX 4090 results.

START WITH YOUR COMPUTER

How much memory do you need?

These scenarios use Ollama, 4-bit model files, 8,192 tokens of context and one chat session. NVIDIA estimates assume 64 GiB of system RAM and no attached display. Mac estimates reserve at least 6 GiB for macOS. Open a setup to change those assumptions or compare runtimes.

Qwen3.8-27B · calculated memory at 8K context
Your hardwareModel filesEstimated peak / availableFit / runtime review
16GB NVIDIA GPUCheck this setup →Q4_K_MOllama · 15.7 GiB weights17.5 GiB15.5 GiB availabledoes not fitRuntime documented
24GB NVIDIA GPUCheck this setup →Q4_K_MOllama · 15.7 GiB weights17.5 GiB23.5 GiB availablecomfortableRuntime documented
32GB NVIDIA GPUCheck this setup →Q4_K_MOllama · 15.7 GiB weights17.5 GiB31.5 GiB availablecomfortableRuntime documented
32GB Apple SiliconCheck this setup →Q4_K_MOllama · 15.7 GiB weights18.8 GiB26.0 GiB availablecomfortableRuntime documented
64GB Apple SiliconCheck this setup →Q4_K_MOllama · 15.7 GiB weights18.8 GiB56.3 GiB availablecomfortableRuntime documented

Memory estimates are not speed tests or a guarantee of fit. System RAM is not added to dedicated VRAM; partial CPU offload is outside these scenarios. Q5, Q6 and Q8 are not verified for this model in the planner.

Mac setups start with Ollama. MLX LM remains available for comparison, but its reviewed Qwen3.8 implementation does not enforce the requested cache-size limit, so those commands are withheld.

Source checks are dated below. Calculator inputs and recorded benchmarks have separate review dates.

SOURCE HEALTHreview required

Last check: 2026-09-27 03:43 UTC · 3 changes for review · 0 unavailable

Inspect source health →
reportedReviewed Aug 18, 2026

What is established

Parameters
27.78B
Architecture
Qwen3_5ForConditionalGeneration
Transformer blocks
64
Full-attention blocks
16
Native context
262,144 tokens
FP16 KV cache
64 KiB / token
Repository revision
1d4bf0f2ff60…
License
Apache-2.0
EXACT FILE · OBSERVED TEST

Which GGUF did Kingy test?

Qwen3.8-27B-Q4_K_M.gguf

15.9 GiB on disk · unsloth/Qwen3.8-27B-GGUF

This is the file used in the linked llama.cpp RTX 4090 tests. The Ollama and MLX planner artifacts have their own file sizes and identifiers; the test does not measure their speed.

Inspect the pinned GGUF file ↗
Revision and checksum

Revision: f1bfb127c64f7072bdd2cad55f258b9c8b2910fe

SHA-256: 7e78da5d7e3ae28d178121f58646953305f3e5bd3cb46f4a75584e8b6c6fe169

Read the GGUF selection guide →

NEXT EVIDENCE

What still needs to be measured

Other quantizations

Repeat the fit boundary across Q5, Q6 and Q8 using pinned artifacts.

Other hardware

Other GPUs, Macs and CPU-only systems need separate receipts.

Quality loss

A pre-registered evaluation across identical prompts and settings.

PROVENANCE

Primary sources

  1. Qwen/Qwen3.8-27B registry recordHugging Face / Qwen · retrieved Aug 18, 2026
    Open source ↗
  2. Qwen3.8-27B immutable configQwen · retrieved Aug 18, 2026
    Open source ↗
Kingy AI Local Lab

Evidence first. Estimates labelled. Corrections preserved.

JSONCSVMethodologyMore Kingy tools