MODEL × HARDWARERESEARCH BETA

Can I actually
run it?

A compatibility lab for local AI. We separate what fits on paper from what has actually survived a run.

3 owned-hardware results10 retained cloud attemptsreceipt-v2 evidenceFailures preserved
PLANNING TOOL

Start with your computer. Or your next model.

Explore NVIDIA GPUs, DGX Spark, Apple Silicon or system RAM. Compare exact local model files, see the memory arithmetic, and get commands when runtime support is verified.

Find models that fit your computer
CALCULATED EXAMPLE

24 GB NVIDIA GPU

Everyday chat · 8K context · one session

Memory pool
Dedicated VRAM
Result
Ranked exact artifacts
Evidence
Estimates stay labelled
DIRECT HARDWARE ANSWERS

Start with a common memory target

Best local LLMs for 24GB VRAMCalculated fits · Save or download your setup →What runs on M4 Max 64GB?Calculated fits · Save or download your setup →What can DGX Spark run?70B and 122B options · One CUDA unified-memory pool →
OWNED-HARDWARE PILOT

Exactly three measured Local Lab results

Retained runs on one Ryzen 9 7950X / RTX 4090 / 128 GB system. Each card separates synthetic decode, context-fill memory, and streaming power so unlike measurements are not blended.

OWNED HARDWARE · REVIEW PASSObserved

Qwen3.8-27B

Ryzen 9 7950X · RTX 4090 24 GB · 128 GB RAM

Runtime
llama.cpp b10453 · CUDA · full GPU offload
Quantization
Q4_K_M
Context
64K allocated · FP16 KV · passed
Decode
49.09 tok/sMedian of five llama-bench 256-token synthetic decode repetitions
Memory
20,266 MiB (19.79 GiB)Sampled GPU peak during selected 64K context-fill run
Power
65.90 W median · 281.38 W peakMedian and peak across retained 8K streaming telemetry
Failure conditions

No runtime failure was observed in the retained 8K, 16K, 32K, or 64K FP16-KV context-fill runs. Contexts above 64K were not tested in this owned-hardware series.

OWNED HARDWARE · REVIEW PASSObserved

Qwen3.6-27B

Ryzen 9 7950X · RTX 4090 24 GB · 128 GB RAM

Runtime
llama.cpp b10453 · CUDA · full GPU offload
Quantization
Q4_K_M
Context
64K allocated · FP16 KV · passed
Decode
49.04 tok/sMedian of five llama-bench 256-token synthetic decode repetitions
Memory
20,266 MiB (19.79 GiB)Sampled GPU peak during selected 64K context-fill run
Power
64.24 W median · 337.72 W peakMedian and peak across retained 8K streaming telemetry
Failure conditions

No runtime failure was observed in the retained 8K, 16K, 32K, or 64K FP16-KV context-fill runs. Contexts above 64K were not tested in this owned-hardware series.

OWNED HARDWARE · REVIEW PASSObserved

Gemma 4 31B-it

Ryzen 9 7950X · RTX 4090 24 GB · 128 GB RAM

Runtime
llama.cpp b10453 · CUDA · full GPU offload
Quantization
Q4_K_M
Context
64K allocated · Q8_0 KV · passed (FP16 KV failed)
Decode
45.00 tok/sMedian of five llama-bench 256-token synthetic decode repetitions
Memory
21,956 MiB (21.44 GiB)Sampled GPU peak during selected 64K context-fill run
Power
63.79 W median · 333.83 W peakMedian and peak across retained 8K streaming telemetry
Failure conditions

64K with FP16 KV failed during context allocation: cudaMalloc could not allocate a 1,200 MiB KV-cache buffer. The same 64K workload passed after changing only KV cache to Q8_0; 32K FP16 KV also passed.

Gate: this pilot stops at three reviewed cards. Expand only after meaningful search impressions, citations, or subscriber demand.

SOURCE FRESHNESS

What is monitored—and what still needs review.

LAST SOURCE CHECK2026-09-20 22:29 UTCRead-only discovery
LAST BENCHMARK2026-08-19 02:20 UTCReviewed evidence only
UPDATE STATEReview required1 item awaits review

Source discovery checks official model revisions, pinned GGUF artifacts, llama.cpp releases and hardware source pages. It can create review work, but it cannot run benchmarks, spend money, alter an Observed label, deploy or publish.

model revisionQwen3.8-27B official revisionChecked 2026-09-20 22:29 UTC
Current
model revisionQwen3.6-27B official revisionChecked 2026-09-20 22:29 UTC
Current
model revisionQwen3-8B official revisionChecked 2026-09-20 22:29 UTC
Current
gguf artifactqwen3.8-27b Qwen3.8-27B-Q4_K_M.ggufChecked 2026-09-20 22:29 UTC
Current
runtime releasellama.cpp releasesChecked 2026-09-20 22:29 UTC
Review required
hardware specificationNVIDIA GeForce RTX 4090 specificationsChecked 2026-09-20 22:29 UTC
Current
hardware specificationNVIDIA GeForce RTX 5090 specificationsChecked 2026-09-20 22:29 UTC
Current
hardware specificationMac Studio M4 Max · 64GB specificationsChecked 2026-09-20 22:29 UTC
Current
hardware specificationMac Studio M3 Ultra · 512GB specificationsChecked 2026-09-20 22:29 UTC
Current
ObservedRun receipt + raw measurements
VerifiedEvidence independently checked
ReportedNamed external source
EstimatedTransparent formula, not a test
OUR STANDARD

Every number needs a receipt.

Model revision. File hash. Runtime version. Context. KV cache. Peak memory. Prompt and generation speed. Failure boundary. If it cannot be reproduced, it is not an observed result.

  1. 01
    IdentifyImmutable model and hardware records
  2. 02
    RunWarm-up plus repeated measurements
  3. 03
    VerifySchema, hash and consistency checks
  4. 04
    PublishRaw receipt and limitations included
Kingy AI Local Lab

Evidence first. Estimates labelled. Corrections preserved.

JSONCSVMethodologyMore Kingy tools