Kingy AI
Explore
Work with Kingy
Menu
OWNED HARDWARE · INDEPENDENT RECOMPUTATION PASS

Qwen3.8-27B × RTX 4090

Q4_K_M on the retained Ryzen 9 7950X / RTX 4090 / 128 GB system. All displayed measurements were recomputed from raw retained evidence; no physical test was rerun.

Reviewed observed result

64K allocated · FP16 KV · passed

No runtime failure was observed in the retained 8K, 16K, 32K, or 64K FP16-KV context-fill runs. Contexts above 64K were not tested in this owned-hardware series.

Runtime
llama.cpp b10453 · CUDA · full GPU offload
Synthetic decode median
49.0884 tok/s
Selected context peak
20,266 MiB (19.79 GiB)
Streaming power
65.90 W median · 281.38 W peak
MEASUREMENT BOUNDARIES

Three profiles, kept separate

49.0884 tok/s

Median of five llama-bench 256-token synthetic decode repetitions. This is a synthetic decode rate, not end-to-end application throughput.

20,266 MiB (19.79 GiB)

Sampled GPU peak during selected 64K context-fill run. The context-fill profile is distinct from the decode benchmark.

65.90 W median

Median and peak across retained 8K streaming telemetry; peak was 281.38 W.

PINNED CONFIGURATION

The exact observed match

Model
Qwen3.8-27B
Artifact quantization
Q4_K_M
Allocated context
65,536 tokens
KV cache
F16
GPU / host
Ryzen 9 7950X · RTX 4090 24 GB · 128 GB RAM
FAILURE CONDITIONS

What failed—and what was not tested

No runtime failure was observed in the retained 8K, 16K, 32K, or 64K FP16-KV context-fill runs. Contexts above 64K were not tested in this owned-hardware series.

These measurements apply only to the pinned GGUF, llama.cpp build, runtime flags, KV precision, and this one physical system. They are not universal RTX 4090 figures.

Kingy AI Local Lab

Evidence first. Estimates labelled. Corrections preserved.

JSONCSVMethodologyMore Kingy tools