Median of five llama-bench 256-token synthetic decode repetitions. This is a synthetic decode rate, not end-to-end application throughput.
Qwen3.6-27B × RTX 4090
Q4_K_M on the retained Ryzen 9 7950X / RTX 4090 / 128 GB system. All displayed measurements were recomputed from raw retained evidence; no physical test was rerun.
64K allocated · FP16 KV · passed
No runtime failure was observed in the retained 8K, 16K, 32K, or 64K FP16-KV context-fill runs. Contexts above 64K were not tested in this owned-hardware series.
- Runtime
- llama.cpp b10453 · CUDA · full GPU offload
- Synthetic decode median
- 49.0424 tok/s
- Selected context peak
- 20,266 MiB (19.79 GiB)
- Streaming power
- 64.24 W median · 337.72 W peak
Three profiles, kept separate
Sampled GPU peak during selected 64K context-fill run. The context-fill profile is distinct from the decode benchmark.
Median and peak across retained 8K streaming telemetry; peak was 337.72 W.
The exact observed match
- Model
- Qwen3.6-27B
- Artifact quantization
- Q4_K_M
- Allocated context
- 65,536 tokens
- KV cache
- F16
- GPU / host
- Ryzen 9 7950X · RTX 4090 24 GB · 128 GB RAM
- Runtime commit
3cb7ffb1a1f612d5e4a46244ae5a3c77ad934a70- Review status
- pass
- Receipt-v2 package
- Package README ↗
- Independent review
- JSON attestation ↗
What failed—and what was not tested
No runtime failure was observed in the retained 8K, 16K, 32K, or 64K FP16-KV context-fill runs. Contexts above 64K were not tested in this owned-hardware series.
These measurements apply only to the pinned GGUF, llama.cpp build, runtime flags, KV precision, and this one physical system. They are not universal RTX 4090 figures.