Median of five llama-bench 256-token synthetic decode repetitions. This is a synthetic decode rate, not end-to-end application throughput.
Gemma 4 31B-it × RTX 4090
Q4_K_M on the retained Ryzen 9 7950X / RTX 4090 / 128 GB system. All displayed measurements were recomputed from raw retained evidence; no physical test was rerun.
64K allocated · Q8_0 KV · passed (FP16 KV failed)
64K with FP16 KV failed during context allocation: cudaMalloc could not allocate a 1,200 MiB KV-cache buffer. The same 64K workload passed after changing only KV cache to Q8_0; 32K FP16 KV also passed.
- Runtime
- llama.cpp b10453 · CUDA · full GPU offload
- Synthetic decode median
- 44.9998 tok/s
- Selected context peak
- 21,956 MiB (21.44 GiB)
- Streaming power
- 63.79 W median · 333.83 W peak
Three profiles, kept separate
Sampled GPU peak during selected 64K context-fill run. The context-fill profile is distinct from the decode benchmark.
Median and peak across retained 8K streaming telemetry; peak was 333.83 W.
The exact observed match
- Model
- Gemma 4 31B-it
- Artifact quantization
- Q4_K_M
- Allocated context
- 65,536 tokens
- KV cache
- Q8_0
- GPU / host
- Ryzen 9 7950X · RTX 4090 24 GB · 128 GB RAM
- Runtime commit
3cb7ffb1a1f612d5e4a46244ae5a3c77ad934a70- Review status
- pass
- Receipt-v2 package
- Package README ↗
- Independent review
- JSON attestation ↗
What failed—and what was not tested
64K with FP16 KV failed during context allocation: cudaMalloc could not allocate a 1,200 MiB KV-cache buffer. The same 64K workload passed after changing only KV cache to Q8_0; 32K FP16 KV also passed.
These measurements apply only to the pinned GGUF, llama.cpp build, runtime flags, KV precision, and this one physical system. They are not universal RTX 4090 figures.