PERMANENT RESULT · OBSERVED

Qwen3.8-27B × RTX 4090

Q4_K_M · FP16 KV cache · llama.cpp · one 24GB RTX 4090. Reviewed receipts establish an allocated-context VRAM-fit boundary, not a full-length prompt result.

observed

Largest successful allocation: 124,928 tokens.

The next 1,024-token allocation, 125,952, failed with CUDA out of memory. Successful measurements used 512 prompt tokens plus 128 generated tokens; they do not prove that a near-124K prompt was processed.

32K peak
17.7 GiB
32K prompt median
2006.8 tok/s
32K generation median
47.7 tok/s
32K TTFT median
262 ms
FIXED ALLOCATIONS

Three measured short requests after one warm-up

Allocated contextOutcomePeak GPUPrompt medianGeneration medianTTFT medianReceipt
8KSuccess16.2 GiB2036.0 tok/s47.7 tok/s259 msJSON ↗
32KSuccess17.7 GiB2006.8 tok/s47.7 tok/s262 msJSON ↗
64KSuccess19.8 GiB2016.0 tok/s47.7 tok/s261 msJSON ↗
128KCUDA OOM15.4 GiB sampledJSON ↗

Every successful row used a 512-token prompt and generated 128 tokens in each measured repetition. The allocated context determines KV-cache capacity; it is not the number of prompt tokens exercised. Failure rows report no synthetic speed values, and pre-failure memory samples are labelled rather than presented as completed peaks.

VRAM-FIT BOUNDARY

A measured allocation bracket, not a long-prompt claim

124,928 allocated · success

23.5 GiB peak. One warm-up and three 512-input/128-output requests completed.

Open success receipt →
125,952 allocated · CUDA OOM

23.5 GiB sampled before failure. The warm-up attempt triggered the preserved CUDA error.

Open failure receipt →
Resolution · 1,024 tokens

The allocation limit is bracketed between these steps. A near-full-context prompt was not tested.

REPRODUCIBILITY

The exact thing that was tested

Artifact

File
Qwen3.8-27B-Q4_K_M.gguf
Size
17,106,775,008 bytes
SHA-256
7e78da5d7e3ae28d178121f58646953305f3e5bd3cb46f4a75584e8b6c6fe169
GGUF revision
f1bfb127c64f7072bdd2cad55f258b9c8b2910fe
Upstream base pin
1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0

The base pin is a recorded upstream reference. The GGUF conversion lineage to that exact commit was not independently established.

Runtime and machine

llama.cpp commit
6d05498314db1b57f81c271080018aa2d0b89be9
GPU
NVIDIA GeForce RTX 4090 · 24,564 MiB
Driver / toolkit
580.65.06 / CUDA 12.8.93
KV / offload
FP16 K+V · all layers on GPU
Provider
RunPod Secure Cloud · EUR-IS-1

Download every raw log and receipt →

Raw archive SHA-256: 028781080144a73a611a6b729edf1af8e8fd4173bd047ec3a3dfc4bab86359b9

LIMITS

What this result does not prove

This does not prove that a prompt near 124,928 tokens can be processed: successful measured requests contained 512 prompt tokens and 128 generated tokens. RunPod is not a retail desktop, so CPU allocation, thermals and power behavior can differ. llama.cpp reported that the GGUF artifact’s extra blk.64/MTP tensors were unused. Quality loss was not evaluated. The allocation boundary must not be generalized to a different GGUF, llama.cpp commit, KV precision, batch size, GPU or layer-offload policy.

EMBED THIS RESULT<iframe src="https://lab.kingy.ai/embed/qwen3.8-27b/rtx-4090-24gb" title="Qwen3.8-27B RTX 4090 compatibility" loading="lazy"></iframe>