OWNED HARDWARE · INDEPENDENT RECOMPUTATION PASS

Gemma 4 31B-it × RTX 4090

Q4_K_M on the retained Ryzen 9 7950X / RTX 4090 / 128 GB system. All displayed measurements were recomputed from raw retained evidence; no physical test was rerun.

Reviewed observed result

64K allocated · Q8_0 KV · passed (FP16 KV failed)

64K with FP16 KV failed during context allocation: cudaMalloc could not allocate a 1,200 MiB KV-cache buffer. The same 64K workload passed after changing only KV cache to Q8_0; 32K FP16 KV also passed.

Runtime
llama.cpp b10453 · CUDA · full GPU offload
Synthetic decode median
44.9998 tok/s
Selected context peak
21,956 MiB (21.44 GiB)
Streaming power
63.79 W median · 333.83 W peak
MEASUREMENT BOUNDARIES

Three profiles, kept separate

44.9998 tok/s

Median of five llama-bench 256-token synthetic decode repetitions. This is a synthetic decode rate, not end-to-end application throughput.

21,956 MiB (21.44 GiB)

Sampled GPU peak during selected 64K context-fill run. The context-fill profile is distinct from the decode benchmark.

63.79 W median

Median and peak across retained 8K streaming telemetry; peak was 333.83 W.

PINNED CONFIGURATION

The exact observed match

Model
Gemma 4 31B-it
Artifact quantization
Q4_K_M
Allocated context
65,536 tokens
KV cache
Q8_0
GPU / host
Ryzen 9 7950X · RTX 4090 24 GB · 128 GB RAM
FAILURE CONDITIONS

What failed—and what was not tested

64K with FP16 KV failed during context allocation: cudaMalloc could not allocate a 1,200 MiB KV-cache buffer. The same 64K workload passed after changing only KV cache to Q8_0; 32K FP16 KV also passed.

These measurements apply only to the pinned GGUF, llama.cpp build, runtime flags, KV precision, and this one physical system. They are not universal RTX 4090 figures.