OWNED HARDWARE · INDEPENDENT RECOMPUTATION PASS

Qwen3.6-27B × RTX 4090

Q4_K_M on the retained Ryzen 9 7950X / RTX 4090 / 128 GB system. All displayed measurements were recomputed from raw retained evidence; no physical test was rerun.

Reviewed observed result

64K allocated · FP16 KV · passed

No runtime failure was observed in the retained 8K, 16K, 32K, or 64K FP16-KV context-fill runs. Contexts above 64K were not tested in this owned-hardware series.

Runtime
llama.cpp b10453 · CUDA · full GPU offload
Synthetic decode median
49.0424 tok/s
Selected context peak
20,266 MiB (19.79 GiB)
Streaming power
64.24 W median · 337.72 W peak
MEASUREMENT BOUNDARIES

Three profiles, kept separate

49.0424 tok/s

Median of five llama-bench 256-token synthetic decode repetitions. This is a synthetic decode rate, not end-to-end application throughput.

20,266 MiB (19.79 GiB)

Sampled GPU peak during selected 64K context-fill run. The context-fill profile is distinct from the decode benchmark.

64.24 W median

Median and peak across retained 8K streaming telemetry; peak was 337.72 W.

PINNED CONFIGURATION

The exact observed match

Model
Qwen3.6-27B
Artifact quantization
Q4_K_M
Allocated context
65,536 tokens
KV cache
F16
GPU / host
Ryzen 9 7950X · RTX 4090 24 GB · 128 GB RAM
FAILURE CONDITIONS

What failed—and what was not tested

No runtime failure was observed in the retained 8K, 16K, 32K, or 64K FP16-KV context-fill runs. Contexts above 64K were not tested in this owned-hardware series.

These measurements apply only to the pinned GGUF, llama.cpp build, runtime flags, KV precision, and this one physical system. They are not universal RTX 4090 figures.