Largest successful FP16 KV allocation in the 1,024-token VRAM-fit search.
Can I actually
run it?
A compatibility lab for local AI. We separate what fits on paper from what has actually survived a run.
Check a configuration
Allocation loaded; short request completed.
Qwen3.8-27B · Q4_K_M · 32K allocated context
Reviewed receipt: 17.7 GiB sampled peak and 47.7 tok/s median generation for 512-input/128-output requests; this is not a full-context prompt test.
- Weights
- 16.7 GB
- KV cache
- 2.0 GB
- Runtime
- 1.3 GB–2.2 GB
- Reserve
- 1.4 GB
Important: observed labels apply only when this exact model, hardware, quantization and context matches a reviewed receipt. Every other selection remains a transparent estimate.
Qwen3.8-27B Q4_K_M on one RTX 4090
The immediately tested next allocation failed; raw server and memory logs are retained.
Median generation speed for 512-input/128-output requests after warm-up.
Scope: this is an allocated-context capacity result, not a near-full-context prompt test. Read the result and inspect receipts →
Every number needs a receipt.
Model revision. File hash. Runtime version. Context. KV cache. Peak memory. Prompt and generation speed. Failure boundary. If it cannot be reproduced, it is not an observed result.
- 01IdentifyImmutable model and hardware records
- 02RunWarm-up plus repeated measurements
- 03VerifySchema, hash and consistency checks
- 04PublishRaw receipt and limitations included