Two DGX Sparks · GB10 · 128 GB each

Spark Swarm Lab

Running notebook of how fast two NVIDIA DGX Sparks can serve many AI agents at once, all sharing one long system prompt and memory file. Every number on this page is measured on the hardware and regenerated from the raw result files.

Updated 6 October 2026, 21:20 · 57 benchmark points · 15 serving configs · 32 graded agent runs
Best on one Spark267tok/s 64 agents sharing a 54K-token context, 27B NVFP4 + MTP2 + shared-prefix checkpoint + FP8 KV, 64 seqs
Graded agent swarm16/16perfect 16 agents at once on the hidden-test reasoning task, all done in about an hour
Fastest context load2,906tok/s Flash-Next over both Sparks; 54K tokens in under 19 s
One agent, Flash-Next39tok/sSplit over both Sparks with MTP3
Spark-to-Spark link172Gb/s NCCL all-gather over ConnectX-7 RoCE, both rails
Scaling

Throughput as agents pile on

Each agent shares a 54,282-token prefix (system prompt plus memory file from a real shop codebase) that is loaded into the cache once, then asks its own question. Hover a point for per-agent speed and time to first token.

tok/s combined05010015020025012481632Agents running at once (log scale)MTP1 + shared-prefix ckptMTP2 + shared-prefix ckptNVFP4, FP8 KVFP8 (0.29)Flash-Next, 2 SparksveloGB10 Flash-Next
  • 27B NVFP4 + MTP1 + shared-prefix checkpoint · vLLM 0.31
  • 27B NVFP4 + MTP2 + shared-prefix checkpoint · vLLM 0.31
  • 27B NVFP4 + FP8 KV cache · vLLM 0.31
  • 27B FP8 · vLLM 0.29
  • Flash-Next NVFP4, split over both Sparks + MTP3 · vLLM 0.31 · both Sparks
  • Flash-Next EXL3 3.05 bpw, one Spark · veloGB10 0.7.3
Config sweep

Every serving config tried

Combined output tokens per second on one Spark unless marked. Sorted by the 32-agent result; the best single-Spark figure in each column is highlighted. Gate 0 checks that a cache hit returns exactly the same answer as a fresh run.

Serving config1 agent83264First token at 32 (s)Cached at 32Gate 0
27B NVFP4 + MTP1 + shared-prefix checkpoint
vLLM 0.31
14.190.1203.4–5.199.2%✓ passed
27B NVFP4 + MTP2 + shared-prefix checkpoint + FP8 KV, 64 seqs
vLLM 0.31
15.980.5188.7267.216.897.2%✓ passed
27B NVFP4 + MTP2 + shared-prefix checkpoint
vLLM 0.31 · Checkpoints state where agents' prompts diverge
16.380.6180.0–12.997.2%✓ passed
27B NVFP4, 0.8 memory, 64 sequences
vLLM 0.31
9.758.4118.8–24.498.2%✓ passed
27B NVFP4 + FP8 KV cache
vLLM 0.31 · About 2x KV capacity
9.858.8116.5–21.898.2%✓ passed
27B NVFP4
vLLM 0.31
10.057.9115.5–23.298.2%✓ passed
27B FP8
vLLM 0.29 · No speculation
7.346.2108.8–6.399.6%✓ passed
27B FP8
vLLM 0.31
6.744.5108.1–5.799.6%✓ passed
Flash-Next NVFP4, split over both Sparks + MTP3
vLLM 0.31 · both Sparks · Tensor parallel over the 200G link
29.4121.1103.7–32.694.3%varies run to run
27B NVFP4 + MTP1
vLLM 0.31
14.462.6101.2–39.596.3%✓ passed
27B NVFP4 + MTP2 + async scheduling
vLLM 0.31
16.060.684.0–59.694.3%✓ passed
27B NVFP4 + MTP2 + FP8 KV cache
vLLM 0.31
15.459.183.9–59.894.3%✓ passed
27B NVFP4 + MTP2
vLLM 0.29 · Option B as first run
17.460.483.5–62.594.3%✓ passed
27B NVFP4 + MTP2
vLLM 0.31
16.459.583.2–60.294.3%✓ passed
Flash-Next EXL3 3.05 bpw, one Spark
veloGB10 0.7.3 · Per-lane cache: the shared context is not reused across agents
49.917.9––––✓ passed
Real work

Graded agent swarms

Codex and OpenCode agents each implement a parcel-carrier selector from a one-page spec. 51 hidden tests grade every run; the same task, prompt and grader used on the M4 Max. Eight agents per Spark, all at the same time.

RoundAgentPerfect (51/51)Per runWhole batchHit 90-min cap
Round 1 · 27B FP8OpenCode3 / 857–90 min90 min1
Round 1 · 27B FP8Codex8 / 869–90 min90 min5
Round 2 · 27B NVFP4 + MTP2, thinking budgetCodex8 / 826–52 min52 min–
Round 2 · 27B NVFP4 + MTP2, thinking budgetOpenCode8 / 827–60 min60 min–
Recipe

Current best single-Spark config

Qwen3.8-27B NVFP4 with one-token MTP speculation and the shared-prefix checkpoint. Gate 0 passed.

docker run -d --name vllm --gpus all --ipc=host --network host \
  -v ~/.cache/huggingface:/root/.cache/huggingface -v ~/.cache/vllm:/root/.cache/vllm \
  -e VLLM_USE_DEEP_GEMM=0 vllm/vllm-openai:v0.31.0 \
  unsloth/Qwen3.8-27B-NVFP4 --served-model-name qwen-local --port 8000 \
  --gpu-memory-utilization 0.7 --max-model-len 131072 --max-num-seqs 32 \
  --max-num-batched-tokens 8192 \
  --enable-prefix-caching --mamba-cache-mode align --enable-prompt-tokens-details \
  --enable-mamba-shared-prefix-checkpoint --prefix-match-unit 16 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":1}' \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning-parser qwen3
Reference

The M4 Max on the same models

MacBook Pro M4 Max, 128 GB, oMLX 0.7.0, measured 27 Sep – 4 Oct 2026. One request at a time unless noted.

M4 Max, 128 GB (oMLX)1 agent, short prompt1 agent after 28K promptPrompt readingSeveral at once
Qwen3.8-27B 8-bit25-28 tok/s14 tok/s~165 tok/sNot measured
Qwen3.8 Flash-Next 4-bit50-57 tok/s51 tok/s~580 tok/s61 tok/s (2 at once)
Notebook

Experiment log

  1. 267 tok/s with 64 agents on one Spark

    Stacking the shared-prefix checkpoint with MTP2, an FP8 KV cache and 64 concurrent sequences at 0.8 memory: 188.7 tok/s at 32 agents and 267.2 tok/s at 64, with the median first token at 11 s. Throughput was still rising at 64 agents. Both Sparks running this would serve about 534 tok/s across 128 agents.

  2. veloGB10: fastest single agent, but no shared cache between agents

    veloGB10 0.7.3, a Rust + CUDA engine written for GB10, runs Qwen3.8-Flash-Next (EXL3 3.05 bpw, 85 GB) on one Spark, which vLLM cannot. Boots in 90 s. One agent: 52.5 tok/s with the 54K context fully cached and a 0.39 s first token, deterministic output, token-exact prefix reuse. Every additional agent re-read the whole 54K context from scratch, because its prefix checkpoints are kept per lane: 16 agents managed only 19 tok/s combined and the slowest waited 277 s for a first token. Best engine tested for one fast lead agent; not suited to a shared-context swarm yet.

  3. Lighter speculation wins: 203 tok/s with MTP1

    The same shared-prefix checkpoint with one-token MTP instead of two reached 203 tok/s at 32 agents and 90 tok/s at 8, with the first token in 5.1 s. Drafting fewer tokens wastes less compute when 32 agents share one GPU, and the checkpoint keeps every agent on the cached 54K prefix. Gate 0 passed.

  4. Shared-prefix checkpoint: 180 tok/s on one Spark

    vLLM 0.31's --enable-mamba-shared-prefix-checkpoint saves state at the point where all agents' prompts diverge. With MTP2 it lifted 32-agent throughput from 83 to 180 tok/s (+116%) and cut time-to-first-token from 60 s to 13 s. Best single-Spark result so far; both Sparks together would serve about 360 tok/s across 64 agents.

  5. MTP re-reads the shared context

    With MTP on, vLLM checkpoints the hybrid model's state only at each prompt's end, so every agent re-processed 1-3K tokens of the shared prefix. At 32 agents that queued about 100K tokens of repeated prefill, pushing time-to-first-token to 60 s and capping throughput at 83 tok/s.

  6. Flash-Next across both Sparks: 2,905 tok/s prefill

    Qwen3.8-Flash-Next NVFP4 split across both Sparks over the ConnectX-7 link on stock vLLM 0.31. 39 tok/s for a single agent, 121 tok/s combined at 8 agents, and the 54K shared context loaded in 18.7 s (about 18x the M4 Max on the 27B). Output varies run to run at temperature 0 even without caching, a known property of this model's sparse-attention kernel on GB10.

  7. Swarm round 2: 16 of 16 perfect in under an hour

    27B NVFP4 with MTP speculative decoding, one model per Spark. 8 Codex agents and 8 OpenCode agents ran at once; every run scored 51/51 on the hidden tests. Batches finished in 52 and 60 minutes. The M4 Max would need about 2.5-3 hours to run the same 16 tasks one at a time.

  8. Swarm round 1: OpenCode agents ran out of thinking room

    8 Codex agents scored 8/8 perfect on the graded reasoning task but took 69-90 minutes each. 5 of 8 OpenCode agents scored 0: the model reasoned for 16,382 tokens, hit OpenCode's output cap and never wrote code. Fixed with a per-step thinking budget of 8,192 tokens and a 32K output cap.

  9. 27B FP8 baseline: 109 tok/s with 32 agents on one Spark

    Qwen3.8-27B FP8 on vLLM 0.29. 32 agents sharing a 54K-token context reached 109 tok/s combined, 14.5x a single agent, with 99.6% of every prompt served from cache. The shared context loaded at about 830 tok/s, roughly 5x the M4 Max on the same model.

  10. Gate 0: no config is benchmarked until it proves its cache is correct

    Every serving config must pass three checks first: a prefix-cache hit on a 20K-token prompt, an identical answer cold and warm at temperature 0 (guards against a known GB10 bug that zeroes hybrid-model state on cache hits), and a correctly formed tool call.

  11. Remote access and Ollama on the tailnet

    Both Sparks joined the tailnet. Ollama on spark01 is exposed tailnet-only through tailscale serve while Ollama itself stays on loopback. gemma3:4b runs fully on the GPU at 79 tok/s.

  12. Cluster link: 172 Gb/s NCCL all-gather

    NCCL 2.30.7 built from source for sm_121. Two-node all_gather at 16 GB: 21.46 GB/s bus bandwidth (about 172 Gb/s), zero errors, traffic spread over both NET/IB rails. That is about 90% of NVIDIA's published raw RDMA figure for this link.

  13. ConnectX-7 is power-gated without a cable

    The 200G NICs were absent from lspci on both boxes. DGX OS powers the ConnectX-7 down when no QSFP cable is attached (saves about 18 W). It hot-appeared the moment the cable went in, with no reboot. One physical port shows up as two interfaces, one per PCIe half, so each gets its own subnet.

  14. Two Sparks on the network

    Both DGX Sparks reached over SSH: DGX OS 7.5.0, kernel 6.17, driver 580.159, CUDA 13.0, 128 GB unified memory and 4 TB NVMe each. Key-based login and short aliases set up. Updated both to DGX OS 7.6.0, kernel 7.0, driver 580.178. Firmware already current.