RUNTIME COMPARISON · SCOPE LOCK

Same model does not mean same test.

Runtime choice changes allocation, supported quantizations, caching and throughput. Results are only compared when the artifact and workload are genuinely compatible.

PERFORMANCE · UNKNOWN

llama.cpp

GGUF · CPU/GPU offload

Receipt must includeVersion, build flags, backend, ngl and cache types
PERFORMANCE · UNKNOWN

MLX

Apple Silicon unified memory

Receipt must includemlx-lm version, chip, macOS and wired-memory state
PERFORMANCE · UNKNOWN

vLLM

GPU serving and concurrency

Receipt must includeVersion, attention backend, utilization cap and tensor parallelism
PERFORMANCE · UNKNOWN

Ollama

Packaged local workflows

Receipt must includeVersion, model manifest, backend and generated runtime settings
COMPARABILITY RULE

No cross-runtime score mixing

Runtime comparisons require the same model revision, equivalent artifact precision, context, batch, concurrency, prompts, output length, warm-up policy and hardware state. Otherwise the site shows separate records and an incompatibility warning.