NVIDIA announced DGX Spark 64GB on October 2, 2026, with availability scheduled for October 23 through Acer, ASUS, Dell, Gigabyte, HP and MSI, starting at US$4,999. For local-AI developers, this adds a memory-capacity choice—not a product confirmed in stock today.
The purchase decision starts with the model, context length and simultaneous requests you need to support. A 64GB system, one 128GB system and a pair of 64GB devices solve different capacity and operational problems.
Hardware timing and software readiness are separate
NVIDIA’s product specifications list GB10 Grace Blackwell, 64GB or 128GB unified memory, ConnectX-7 networking and DGX OS. The 64GB option is OEM-only. The platform architecture shares physical memory between CPU and GPU and uses Arm64 with Ubuntu-based DGX OS; check your runtime’s architecture support, not just its CUDA label.
Documented now: Sync Cluster Assistant configures the inter-device network and SSH. It does not install inference or fine-tuning workloads.
Planned for late October: Sync Model Launcher, including Qwen3.8-27B launching and OpenCode setup.
Still forthcoming: the Blender installer and the announced 64GB-specific vLLM, local-agent and distributed-workload playbooks.
If you need the launcher to operate the system, wait for its release and supported-model documentation. Hardware availability alone will not settle that requirement.
Which configuration should you evaluate?
These are architectural recommendations, not tested product rankings. NVIDIA advertises support for models up to 100 billion parameters on 64GB and 200 billion on 128GB, but those examples do not specify your usable context, concurrency or memory representation.
Choice | A reasonable fit when… | Check before ordering |
|---|---|---|
One 64GB system | The complete workload fits one device and the October schedule suits your project. | Model and cache precision, peak memory, Arm64 runtime support, exact OEM configuration and delivery date. |
One 128GB system | You need more single-node memory without operating distributed inference. | Current price and stock, workload support and sufficient headroom. Extra memory alone does not establish faster inference. |
Two 64GB systems | A supported sharded workload needs both devices, or separate replicas serve useful independent requests. | Actual model/runtime sharding support, cabling and two-host operations. Each replica still needs to fit one device. |
A memory calculation that parameter counts miss
Model weights are only one allocation. Budget for attention cache and recurrent state, runtime buffers, the operating system, companion applications and reserve. Longer conversations and concurrent requests can change the answer even when the checkpoint stays the same.
Consider Qwen3.8-27B as an illustration, not a reconstruction of NVIDIA’s benchmark. The official repository metadata reports 27,781,427,952 BF16 parameters. Assuming all stored weights are resident at two bytes each gives 55.56 decimal GB, or about 51.75 GiB.
Its pinned configuration has 16 full-attention layers, four key/value heads and 256 dimensions per head. Assuming one sequence with conventional, uncompressed two-byte keys and values:
Full-attention cache per token:
16 layers × 4 KV heads × 256 dimensions × 2 (K and V) × 2 bytes
= 65,536 bytes = 64 KiB
At 262,144 cached tokens: 16 GiB
Raw BF16 weights + this cache: about 67.75 GiBThat subtotal exceeds even 64 GiB, before recurrent state, vision processing, activations, temporary buffers, OS and other services. GB and GiB are different units; the calculation uses GiB for the combined total.
This does not mean Qwen3.8-27B cannot run on 64GB. Quantized weights or cache, shorter context and different memory-residency policies change the footprint. It means a model’s maximum context is not a promise that every hardware configuration can use it. No model was run for this calculation.
What two connected systems actually provide
With two connected systems, memory remains on separate hosts. A distributed runtime must divide the model or work between them. Running independent replicas is another option, but duplicates weights instead of enlarging one model’s available memory. Do not treat the nominal 128GB aggregate as one transparent allocation.
The current Cluster Assistant guide supports two or three directly cabled devices, or two to four through a switch; four require a switch. Two is the setup highlighted in the announcement, not the platform-wide maximum. Supported software, physical cabling, devices added to Sync, and SSH/sudo access remain prerequisites.
Crucially, the guide separates network configuration from workload setup. A successful network check does not establish that your application can shard its model across machines. Require a supported recipe for the actual model, quantization and parallelism mode.
The external link is rated at 200 Gbit/s: a nominal 25 GB/s before protocol overhead. The local memory specification is 273 GB/s. These are different data paths, not a measured inference-speed ratio. Their distinction explains why adding a host requires a workload-specific plan.
How to read NVIDIA’s 1.7× claim
NVIDIA reports up to 1.7× performance for two 64GB systems versus a single system in its Qwen3.8-27B test. The inspected announcement does not specify enough about the metric, baseline memory, quantization, context, concurrency or runtime versions to reproduce that comparison.
Treat it as a vendor result, not a promised token rate, latency reduction or comparison with your chosen 128GB SKU. It also does not establish three- or four-device scaling.
What to ask for before buying
Get an itemized OEM quote: memory, storage, warranty, bundled cables, taxes and expected delivery. The announced starting price is not a regional stock or fulfillment guarantee.
Request workload evidence with a fixed checkpoint, quantization, runtime version, context and concurrent-request count. Ask for peak memory, first-token latency, sustained generation and task success—not only whether the model loads.
For two devices, establish whether you need one sharded model or independent replicas. Ask who supports the distributed configuration and which workload steps remain manual.
A simple budget illustration: two units at the announced US$4,999 starting price total US$9,998 before extras. This is arithmetic, not a bundle quote. A current, comparable 128GB price was not verified, so neither the 64GB option nor the two-device route is established here as cheaper.
Local inference also needs a precise privacy boundary. ASUS’s product-page qualifications note that external services and network-enabled integrations can involve separate fees or data exposure. If offline operation matters, assess model execution, agent tools, logs and update/download paths separately; local hardware alone cannot guarantee it.
For a deeper look at weight representation and cache budgeting, see GGUF in Transformers: A Better Test Bench for Local AI. Its Apple Silicon implementation is not a DGX Spark compatibility recommendation.
The practical choice is conditional: evaluate 64GB for a workload that fits, 128GB for more single-node headroom, and two devices for a specifically supported distributed workload or useful replicas. Wait if the software you need is still only announced.
Methodology: AI-assisted reporting and analysis based on NVIDIA, ASUS and Qwen’s published material, checked October 2, 2026. Memory figures are explicit calculations, not measurements. No hardware testing, model execution or independent performance verification was performed.
