Running notebook of how fast two NVIDIA DGX Sparks can serve many AI agents at once, all sharing one long system prompt and memory file. Every number on this page is measured on the hardware and regenerated from the raw result files.
Updated 6 October 2026, 21:20 · 57 benchmark points · 15 serving configs · 32 graded agent runsEach agent shares a 54,282-token prefix (system prompt plus memory file from a real shop codebase) that is loaded into the cache once, then asks its own question. Hover a point for per-agent speed and time to first token.
Combined output tokens per second on one Spark unless marked. Sorted by the 32-agent result; the best single-Spark figure in each column is highlighted. Gate 0 checks that a cache hit returns exactly the same answer as a fresh run.
| Serving config | 1 agent | 8 | 32 | 64 | First token at 32 (s) | Cached at 32 | Gate 0 |
|---|---|---|---|---|---|---|---|
27B NVFP4 + MTP1 + shared-prefix checkpoint vLLM 0.31 | 14.1 | 90.1 | 203.4 | – | 5.1 | 99.2% | ✓ passed |
27B NVFP4 + MTP2 + shared-prefix checkpoint + FP8 KV, 64 seqs vLLM 0.31 | 15.9 | 80.5 | 188.7 | 267.2 | 16.8 | 97.2% | ✓ passed |
27B NVFP4 + MTP2 + shared-prefix checkpoint vLLM 0.31 · Checkpoints state where agents' prompts diverge | 16.3 | 80.6 | 180.0 | – | 12.9 | 97.2% | ✓ passed |
27B NVFP4, 0.8 memory, 64 sequences vLLM 0.31 | 9.7 | 58.4 | 118.8 | – | 24.4 | 98.2% | ✓ passed |
27B NVFP4 + FP8 KV cache vLLM 0.31 · About 2x KV capacity | 9.8 | 58.8 | 116.5 | – | 21.8 | 98.2% | ✓ passed |
27B NVFP4 vLLM 0.31 | 10.0 | 57.9 | 115.5 | – | 23.2 | 98.2% | ✓ passed |
27B FP8 vLLM 0.29 · No speculation | 7.3 | 46.2 | 108.8 | – | 6.3 | 99.6% | ✓ passed |
27B FP8 vLLM 0.31 | 6.7 | 44.5 | 108.1 | – | 5.7 | 99.6% | ✓ passed |
Flash-Next NVFP4, split over both Sparks + MTP3 vLLM 0.31 · both Sparks · Tensor parallel over the 200G link | 29.4 | 121.1 | 103.7 | – | 32.6 | 94.3% | varies run to run |
27B NVFP4 + MTP1 vLLM 0.31 | 14.4 | 62.6 | 101.2 | – | 39.5 | 96.3% | ✓ passed |
27B NVFP4 + MTP2 + async scheduling vLLM 0.31 | 16.0 | 60.6 | 84.0 | – | 59.6 | 94.3% | ✓ passed |
27B NVFP4 + MTP2 + FP8 KV cache vLLM 0.31 | 15.4 | 59.1 | 83.9 | – | 59.8 | 94.3% | ✓ passed |
27B NVFP4 + MTP2 vLLM 0.29 · Option B as first run | 17.4 | 60.4 | 83.5 | – | 62.5 | 94.3% | ✓ passed |
27B NVFP4 + MTP2 vLLM 0.31 | 16.4 | 59.5 | 83.2 | – | 60.2 | 94.3% | ✓ passed |
Flash-Next EXL3 3.05 bpw, one Spark veloGB10 0.7.3 · Per-lane cache: the shared context is not reused across agents | 49.9 | 17.9 | – | – | – | – | ✓ passed |
Codex and OpenCode agents each implement a parcel-carrier selector from a one-page spec. 51 hidden tests grade every run; the same task, prompt and grader used on the M4 Max. Eight agents per Spark, all at the same time.
| Round | Agent | Perfect (51/51) | Per run | Whole batch | Hit 90-min cap |
|---|---|---|---|---|---|
| Round 1 · 27B FP8 | OpenCode | 3 / 8 | 57–90 min | 90 min | 1 |
| Round 1 · 27B FP8 | Codex | 8 / 8 | 69–90 min | 90 min | 5 |
| Round 2 · 27B NVFP4 + MTP2, thinking budget | Codex | 8 / 8 | 26–52 min | 52 min | – |
| Round 2 · 27B NVFP4 + MTP2, thinking budget | OpenCode | 8 / 8 | 27–60 min | 60 min | – |
Qwen3.8-27B NVFP4 with one-token MTP speculation and the shared-prefix checkpoint. Gate 0 passed.
docker run -d --name vllm --gpus all --ipc=host --network host \
-v ~/.cache/huggingface:/root/.cache/huggingface -v ~/.cache/vllm:/root/.cache/vllm \
-e VLLM_USE_DEEP_GEMM=0 vllm/vllm-openai:v0.31.0 \
unsloth/Qwen3.8-27B-NVFP4 --served-model-name qwen-local --port 8000 \
--gpu-memory-utilization 0.7 --max-model-len 131072 --max-num-seqs 32 \
--max-num-batched-tokens 8192 \
--enable-prefix-caching --mamba-cache-mode align --enable-prompt-tokens-details \
--enable-mamba-shared-prefix-checkpoint --prefix-match-unit 16 \
--speculative-config '{"method":"mtp","num_speculative_tokens":1}' \
--enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning-parser qwen3MacBook Pro M4 Max, 128 GB, oMLX 0.7.0, measured 27 Sep – 4 Oct 2026. One request at a time unless noted.
| M4 Max, 128 GB (oMLX) | 1 agent, short prompt | 1 agent after 28K prompt | Prompt reading | Several at once |
|---|---|---|---|---|
| Qwen3.8-27B 8-bit | 25-28 tok/s | 14 tok/s | ~165 tok/s | Not measured |
| Qwen3.8 Flash-Next 4-bit | 50-57 tok/s | 51 tok/s | ~580 tok/s | 61 tok/s (2 at once) |
MTP1 version of the stacked config up to 64 agents, then swarm round 3 on it with all 16 graded agents (8 Codex + 8 OpenCode) on a single Spark. SGLang with DFlash2 after that.
Stacking the shared-prefix checkpoint with MTP2, an FP8 KV cache and 64 concurrent sequences at 0.8 memory: 188.7 tok/s at 32 agents and 267.2 tok/s at 64, with the median first token at 11 s. Throughput was still rising at 64 agents. Both Sparks running this would serve about 534 tok/s across 128 agents.
veloGB10 0.7.3, a Rust + CUDA engine written for GB10, runs Qwen3.8-Flash-Next (EXL3 3.05 bpw, 85 GB) on one Spark, which vLLM cannot. Boots in 90 s. One agent: 52.5 tok/s with the 54K context fully cached and a 0.39 s first token, deterministic output, token-exact prefix reuse. Every additional agent re-read the whole 54K context from scratch, because its prefix checkpoints are kept per lane: 16 agents managed only 19 tok/s combined and the slowest waited 277 s for a first token. Best engine tested for one fast lead agent; not suited to a shared-context swarm yet.
The same shared-prefix checkpoint with one-token MTP instead of two reached 203 tok/s at 32 agents and 90 tok/s at 8, with the first token in 5.1 s. Drafting fewer tokens wastes less compute when 32 agents share one GPU, and the checkpoint keeps every agent on the cached 54K prefix. Gate 0 passed.
vLLM 0.31's --enable-mamba-shared-prefix-checkpoint saves state at the point where all agents' prompts diverge. With MTP2 it lifted 32-agent throughput from 83 to 180 tok/s (+116%) and cut time-to-first-token from 60 s to 13 s. Best single-Spark result so far; both Sparks together would serve about 360 tok/s across 64 agents.
With MTP on, vLLM checkpoints the hybrid model's state only at each prompt's end, so every agent re-processed 1-3K tokens of the shared prefix. At 32 agents that queued about 100K tokens of repeated prefill, pushing time-to-first-token to 60 s and capping throughput at 83 tok/s.
Qwen3.8-Flash-Next NVFP4 split across both Sparks over the ConnectX-7 link on stock vLLM 0.31. 39 tok/s for a single agent, 121 tok/s combined at 8 agents, and the 54K shared context loaded in 18.7 s (about 18x the M4 Max on the 27B). Output varies run to run at temperature 0 even without caching, a known property of this model's sparse-attention kernel on GB10.
27B NVFP4 with MTP speculative decoding, one model per Spark. 8 Codex agents and 8 OpenCode agents ran at once; every run scored 51/51 on the hidden tests. Batches finished in 52 and 60 minutes. The M4 Max would need about 2.5-3 hours to run the same 16 tasks one at a time.
8 Codex agents scored 8/8 perfect on the graded reasoning task but took 69-90 minutes each. 5 of 8 OpenCode agents scored 0: the model reasoned for 16,382 tokens, hit OpenCode's output cap and never wrote code. Fixed with a per-step thinking budget of 8,192 tokens and a 32K output cap.
Qwen3.8-27B FP8 on vLLM 0.29. 32 agents sharing a 54K-token context reached 109 tok/s combined, 14.5x a single agent, with 99.6% of every prompt served from cache. The shared context loaded at about 830 tok/s, roughly 5x the M4 Max on the same model.
Every serving config must pass three checks first: a prefix-cache hit on a 20K-token prompt, an identical answer cold and warm at temperature 0 (guards against a known GB10 bug that zeroes hybrid-model state on cache hits), and a correctly formed tool call.
Both Sparks joined the tailnet. Ollama on spark01 is exposed tailnet-only through tailscale serve while Ollama itself stays on loopback. gemma3:4b runs fully on the GPU at 79 tok/s.
NCCL 2.30.7 built from source for sm_121. Two-node all_gather at 16 GB: 21.46 GB/s bus bandwidth (about 172 Gb/s), zero errors, traffic spread over both NET/IB rails. That is about 90% of NVIDIA's published raw RDMA figure for this link.
The 200G NICs were absent from lspci on both boxes. DGX OS powers the ConnectX-7 down when no QSFP cable is attached (saves about 18 W). It hot-appeared the moment the cable went in, with no reboot. One physical port shows up as two interfaces, one per PCIe half, so each gets its own subnet.
Both DGX Sparks reached over SSH: DGX OS 7.5.0, kernel 6.17, driver 580.159, CUDA 13.0, 128 GB unified memory and 4 TB NVMe each. Key-based login and short aliases set up. Updated both to DGX OS 7.6.0, kernel 7.0, driver 580.178. Firmware already current.