llama.cpp
GGUF · CPU/GPU offload
Receipt must includeVersion, build flags, backend, ngl and cache typesRuntime choice changes allocation, supported quantizations, caching and throughput. Results are only compared when the artifact and workload are genuinely compatible.
GGUF · CPU/GPU offload
Receipt must includeVersion, build flags, backend, ngl and cache typesApple Silicon unified memory
Receipt must includemlx-lm version, chip, macOS and wired-memory stateGPU serving and concurrency
Receipt must includeVersion, attention backend, utilization cap and tensor parallelismPackaged local workflows
Receipt must includeVersion, model manifest, backend and generated runtime settingsRuntime comparisons require the same model revision, equivalent artifact precision, context, batch, concurrency, prompts, output length, warm-up policy and hardware state. Otherwise the site shows separate records and an incompatibility warning.