Kingy AI
Explore
Work with Kingy
Menu
DGX SPARK · CALCULATED TEXT INFERENCE

What can DGX Spark run?

Explore exact 70B and 122B model files on a single Spark. Start with memory fit and runtime support, then inspect the assumptions before downloading.

Source checks are dated below. Calculator inputs and recorded benchmarks have separate review dates.

SOURCE HEALTHreview required

Last check: 2026-09-27 03:43 UTC · 3 changes for review · 0 unavailable

Inspect source health →
THE USEFUL ANSWER

Larger models can fit. Speed remains a separate question.

At 8K context and one concurrent session, the planner evaluates Llama 3.3 70B and Qwen3.5 122B-A10B using their exact Ollama artifacts. The table below is calculated for text inference; Kingy has no measured DGX Spark throughput or latency receipt.

Explore this Spark setup →
Single DGX Spark · 8K context · one text session · Ollama
ModelExact fileCalculated peakHeadroom / fitInspect
Llama 3.3 70B InstructLlama 3.3 Community License39.6 GiB · Q4_K_M42.6 GiB60.6 GiB · comfortableInspect this model →
Qwen3.5 122B-A10BApache-2.075.8 GiB · Q4_K_M76.6 GiB26.6 GiB · comfortableInspect this model →

128GB marketed memory is not 128GiB of free VRAM

NVIDIA specifies a GB10 Grace Blackwell Superchip, a 20-core Arm CPU and 128GB of coherent unified memory. CPU and GPU share this pool; the calculator never adds it twice. The preset interprets the advertised capacity conservatively as 128,000,000,000 bytes (119.2 GiB), then keeps a 16 GiB planning reserve. This reserve is an explicit assumption, not a measured free-memory reading.

DGX Spark uses the CUDA unified-memory adapter. Apple Silicon and discrete NVIDIA GPUs have separate adapters. Actual free memory, display use, OS state and runtime allocations can reduce the budget further. NVIDIA hardware specifications · NVIDIA memory and runtime caveats.

Can the runtime use this hardware?

Ollama documents GB10 GPU support and a Linux Arm64 package; NVIDIA provides a Spark Ollama setup. The calculator additionally checks the selected model’s exact artifact binding and review expiry before showing commands. MLX LM is not supported on Spark. An Arm64 CPU package alone does not prove GPU acceleration for every model. Ollama GPU support · Linux installation · NVIDIA Spark Ollama instructions.

Why the full MoE model still counts

Qwen3.5 122B-A10B activates about 10B parameters per token according to its model card, but that does not reduce resident weights to a 10B file. This planner counts the entire selected artifact, its attention KV cache, recurrent state and runtime overhead. Image inference and its encoder allocations are outside this text-only estimate. Qwen model card.

What this page cannot establish

Memory fit does not guarantee acceptable token speed, latency or task quality. We do not turn advertised FP4 compute or bandwidth into tokens per second. Two Sparks do not become one ordinary memory pool: distributed inference needs a topology, runtime and communication model. Multi-Spark, multi-GPU and CPU/GPU offload are explicitly outside this calculator.

CHANGE ONE DECISION

What if I change…

Compare the same model across alternatives. These are calculated memory results; runtime support is checked separately.

Llama 3.3 70B Instruct · chat · 1 concurrent session
AlternativeArtifactCalculated peakHeadroomMemory fit / runtimeTry it
4-bit artifactExact catalogue files only; precision is not a measured quality score.Q4_K_M42.6 GiB60.6 GiBcomfortableRuntime documentedOpen setup →
5-bit artifactExact catalogue files only; precision is not a measured quality score.Q5_K_M49.5 GiB53.7 GiBcomfortableRuntime documentedOpen setup →
6-bit artifactExact catalogue files only; precision is not a measured quality score.Q6_K56.9 GiB46.3 GiBcomfortableRuntime documentedOpen setup →
8-bit artifactExact catalogue files only; precision is not a measured quality score.Q8_072.8 GiB30.4 GiBcomfortableRuntime documentedOpen setup →

No compatible artifact means unknown here, not that the model can never run. Multiple GPUs, multi-Spark clusters and hybrid offload need a topology-specific calculation.

Keep this setup

Save the settings here, share them, or download a brief with the calculations and sources.

Browser save stays on this device. Shared links restore settings; downloads preserve a dated result snapshot. Estimates are not run receipts.

Compare with 24GB VRAM planning · Explore Apple unified memory · Read the methodology

Kingy AI Local Lab

Evidence first. Estimates labelled. Corrections preserved.

JSONCSVMethodologyMore Kingy tools