23.5 GiB peak. One warm-up and three 512-input/128-output requests completed.
Open success receipt →Qwen3.8-27B × RTX 4090
Q4_K_M · FP16 KV cache · llama.cpp · one 24GB RTX 4090. Reviewed receipts establish an allocated-context VRAM-fit boundary, not a full-length prompt result.
Largest successful allocation: 124,928 tokens.
The next 1,024-token allocation, 125,952, failed with CUDA out of memory. Successful measurements used 512 prompt tokens plus 128 generated tokens; they do not prove that a near-124K prompt was processed.
- 32K peak
- 17.7 GiB
- 32K prompt median
- 2006.8 tok/s
- 32K generation median
- 47.7 tok/s
- 32K TTFT median
- 262 ms
Three measured short requests after one warm-up
| Allocated context | Outcome | Peak GPU | Prompt median | Generation median | TTFT median | Receipt |
|---|---|---|---|---|---|---|
| 8K | Success | 16.2 GiB | 2036.0 tok/s | 47.7 tok/s | 259 ms | JSON ↗ |
| 32K | Success | 17.7 GiB | 2006.8 tok/s | 47.7 tok/s | 262 ms | JSON ↗ |
| 64K | Success | 19.8 GiB | 2016.0 tok/s | 47.7 tok/s | 261 ms | JSON ↗ |
| 128K | CUDA OOM | 15.4 GiB sampled | — | — | — | JSON ↗ |
Every successful row used a 512-token prompt and generated 128 tokens in each measured repetition. The allocated context determines KV-cache capacity; it is not the number of prompt tokens exercised. Failure rows report no synthetic speed values, and pre-failure memory samples are labelled rather than presented as completed peaks.
A measured allocation bracket, not a long-prompt claim
23.5 GiB sampled before failure. The warm-up attempt triggered the preserved CUDA error.
Open failure receipt →The allocation limit is bracketed between these steps. A near-full-context prompt was not tested.
The exact thing that was tested
Artifact
- File
- Qwen3.8-27B-Q4_K_M.gguf
- Size
- 17,106,775,008 bytes
- SHA-256
7e78da5d7e3ae28d178121f58646953305f3e5bd3cb46f4a75584e8b6c6fe169- GGUF revision
f1bfb127c64f7072bdd2cad55f258b9c8b2910fe- Upstream base pin
1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0
The base pin is a recorded upstream reference. The GGUF conversion lineage to that exact commit was not independently established.
Runtime and machine
- llama.cpp commit
6d05498314db1b57f81c271080018aa2d0b89be9- GPU
- NVIDIA GeForce RTX 4090 · 24,564 MiB
- Driver / toolkit
- 580.65.06 / CUDA 12.8.93
- KV / offload
- FP16 K+V · all layers on GPU
- Provider
- RunPod Secure Cloud · EUR-IS-1
Download every raw log and receipt →
Raw archive SHA-256: 028781080144a73a611a6b729edf1af8e8fd4173bd047ec3a3dfc4bab86359b9
What this result does not prove
This does not prove that a prompt near 124,928 tokens can be processed: successful measured requests contained 512 prompt tokens and 128 generated tokens. RunPod is not a retail desktop, so CPU allocation, thermals and power behavior can differ. llama.cpp reported that the GGUF artifact’s extra blk.64/MTP tensors were unused. Quality loss was not evaluated. The allocation boundary must not be generalized to a different GGUF, llama.cpp commit, KV precision, batch size, GPU or layer-offload policy.