MODEL × HARDWARERESEARCH BETA

Can I actually
run it?

A compatibility lab for local AI. We separate what fits on paper from what has actually survived a run.

10 observed attempts7 sourced recordsv0.1 estimatorFailures preserved
01

Check a configuration

Estimator formula v0.1
OBSERVED ATTEMPT

Allocation loaded; short request completed.

Qwen3.8-27B · Q4_K_M · 32K allocated context

Reviewed receipt: 17.7 GiB sampled peak and 47.7 tok/s median generation for 512-input/128-output requests; this is not a full-context prompt test.

ESTIMATED PEAK RANGE21.422.3GB · 24 GB available
Weights
16.7 GB
KV cache
2.0 GB
Runtime
1.3 GB2.2 GB
Reserve
1.4 GB
CONSERVATIVE HEADROOM1.7 GB

Important: observed labels apply only when this exact model, hardware, quantization and context matches a reviewed receipt. Every other selection remains a transparent estimate.

LATEST OBSERVED RESULT

Qwen3.8-27B Q4_K_M on one RTX 4090

124,928 allocated · success

Largest successful FP16 KV allocation in the 1,024-token VRAM-fit search.

125,952 allocated · CUDA OOM

The immediately tested next allocation failed; raw server and memory logs are retained.

47.7 tok/s at 32K allocated

Median generation speed for 512-input/128-output requests after warm-up.

Scope: this is an allocated-context capacity result, not a near-full-context prompt test. Read the result and inspect receipts →

ObservedRun receipt + raw measurements
VerifiedEvidence independently checked
ReportedNamed external source
EstimatedTransparent formula, not a test
OUR STANDARD

Every number needs a receipt.

Model revision. File hash. Runtime version. Context. KV cache. Peak memory. Prompt and generation speed. Failure boundary. If it cannot be reproduced, it is not an observed result.

  1. 01
    IdentifyImmutable model and hardware records
  2. 02
    RunWarm-up plus repeated measurements
  3. 03
    VerifySchema, hash and consistency checks
  4. 04
    PublishRaw receipt and limitations included