24 GB NVIDIA GPU
Everyday chat · 8K context · one session
- Memory pool
- Dedicated VRAM
- Result
- Ranked exact artifacts
- Evidence
- Estimates stay labelled
A compatibility lab for local AI. We separate what fits on paper from what has actually survived a run.
Explore NVIDIA GPUs, DGX Spark, Apple Silicon or system RAM. Compare exact local model files, see the memory arithmetic, and get commands when runtime support is verified.
Find models that fit your computerEveryday chat · 8K context · one session
Retained runs on one Ryzen 9 7950X / RTX 4090 / 128 GB system. Each card separates synthetic decode, context-fill memory, and streaming power so unlike measurements are not blended.
Ryzen 9 7950X · RTX 4090 24 GB · 128 GB RAM
No runtime failure was observed in the retained 8K, 16K, 32K, or 64K FP16-KV context-fill runs. Contexts above 64K were not tested in this owned-hardware series.
Ryzen 9 7950X · RTX 4090 24 GB · 128 GB RAM
No runtime failure was observed in the retained 8K, 16K, 32K, or 64K FP16-KV context-fill runs. Contexts above 64K were not tested in this owned-hardware series.
Ryzen 9 7950X · RTX 4090 24 GB · 128 GB RAM
64K with FP16 KV failed during context allocation: cudaMalloc could not allocate a 1,200 MiB KV-cache buffer. The same 64K workload passed after changing only KV cache to Q8_0; 32K FP16 KV also passed.
Gate: this pilot stops at three reviewed cards. Expand only after meaningful search impressions, citations, or subscriber demand.
Source discovery checks official model revisions, pinned GGUF artifacts, llama.cpp releases and hardware source pages. It can create review work, but it cannot run benchmarks, spend money, alter an Observed label, deploy or publish.
Model revision. File hash. Runtime version. Context. KV cache. Peak memory. Prompt and generation speed. Failure boundary. If it cannot be reproduced, it is not an observed result.
Evidence first. Estimates labelled. Corrections preserved.