vLLM vs SGLang on a DGX Spark: Breaking the 14.5 tok/s Wall
SGLang with the DFlash2 draft model wrote Qwen3.8-27B up to 3.3× faster than vLLM on one DGX Spark. Why the Spark has a speed limit, and what beating it costs.
On one NVIDIA DGX Spark, Qwen3.8-27B writes about 12.6 tokens per second under vLLM, and tuning barely moves it, because the hardware caps it near 14.5. With SGLang and the DFlash2 draft model, the same machine wrote 2 to 3.3 times as fast in our recorded test. A 3D page that took vLLM 464 seconds took SGLang 179.
You can watch both engines race, token by token on QAI EdgeBench.
Why a Spark tops out near 14.5 tokens per second
To write one token, a language model reads all of its weights from memory. The Spark's memory moves 273 GB per second, and Qwen3.8-27B in 4-bit form is about 18.8 GB, so the ceiling is 273 ÷ 18.8, about 14.5 tokens per second (ai-muninn). The engine doesn't change that arithmetic. Without help, community testers measure every engine on the Spark at 8 to 12 tokens per second, and our own vLLM run came in at 12.6.
In an agentic coding session, most of the wait is the model writing: code, edits, tool calls. So writing speed is the lever.
How speculative decoding gets past it
A small draft model guesses several tokens ahead, eight per pass in DFlash2's case. The main model then checks all of them in a single pass, the same pass it would have spent writing one token, and keeps the ones it would have written itself. Each pass now produces several tokens instead of one. Because the main model still approves every token, the answers are as good as without drafting. Only the speed changes.
The gain depends on how many guesses survive, and that depends on the draft model and on how predictable the text is.
| MTP (built in) | DSpark | DFlash2 | |
|---|---|---|---|
| Tokens kept per guess, authors’ test | 4.28 | 3.62 | 4.80 |
| One Spark, one request | 18.5 tok/s (vLLM) | 36.6 tok/s (SGLang) | 47.9 tok/s (SGLang) |
| Our test: tokens kept per pass | not tested | not tested | 3.8–5.0, by scene |
| Known problems | not assessed here | none reported in these sources | Crashes on very short requests and on cancelled replies |
What other Spark owners measured
NVIDIA hasn't published figures for this model on the Spark, so the grey bars below are community results, one request at a time on one Spark. Our own test runs are in blue.
Show as table
| Setup | tok/s | Source |
|---|---|---|
| llama.cpp, 4-bit, no drafting | 11.6 | kubesimplify |
| vLLM, 4-bit, no drafting | 12.6 | QAI EdgeBench, 3D page race |
| vLLM, 4-bit + MTP | 18.5 | OpenZeka study |
| Ollama, 4-bit + MTP, on by default | 26.5 | kubesimplify |
| vLLM, 8-bit + DFlash2 | 31.7 | 0xBakeer |
| SGLang, 4-bit + DSpark | 36.6 | OpenZeka study |
| SGLang, 4-bit + DFlash2 | 41.2 | QAI EdgeBench, 3D page race |
| SGLang, 4-bit + DFlash2 | 47.9 | OpenZeka study |
Without speculative decoding, the engines are about equal, because the memory limit applies to all of them. What separates them is which draft models each one runs well on the Spark.
| Engine | Best reply speed | Setup | Notes |
|---|---|---|---|
| SGLang | 47.9 tok/s | 4-bit + DFlash2 | Fastest. 34–40 on agentic coding. DFlash2 has two known crash bugs. |
| vLLM | 31.7 tok/s | 8-bit + DFlash2 | Measured with 8-bit weights. |
| Ollama | 26.5 tok/s | 4-bit + MTP, on by default | Simplest to run. Slower at reading prompts: 731 tok/s on 2,000-token prompts. |
| llama.cpp | 11.6 tok/s | 4-bit, no drafting | Supports DFlash2, but no Spark measurement was found. 837 tok/s reading 2,000-token prompts. |
| oMLX | n/a | Apple Silicon only | Does not run on the Spark. |
What we measured
We recorded both engines on 3 October 2026, on the same Spark, with the same model size, the same prompts and the same settings. They ran one after the other, because both servers don't fit in memory at once. vLLM ran without speculative decoding; SGLang ran with DFlash2.
Show as table
| vLLM, tok/s | SGLang + DFlash2, tok/s | Difference | Source | |
|---|---|---|---|---|
| 3D galaxy page, race recording | 12.6 | 41.2 | 3.3× faster | Race recording |
| Editing a file, 3,000 tokens | 12.6 | 40.0 | 3.2× faster | Benchmark run, mean of 2 |
| Writing new code, 3,000 tokens | 12.7 | 28.0 | 2.2× faster | Benchmark run, mean of 2 |
| Python tutorial, 512 tokens | 12.7 | 24.9 | 2.0× faster | Decode check |
Editing a file landed at the top of the community's 34 to 40 range, and writing new code fell below it. That fits how drafting works: an edit mostly repeats text the model has just read, so more of the draft survives. The finish times show what that means for waiting.
Show as table
| vLLM, s | SGLang + DFlash2, s | Difference | |
|---|---|---|---|
| 3D galaxy page | 463.9 | 178.8 | 2.6× sooner |
| Editing a file | 60.7 | 19.7 | 3.1× sooner |
| Writing new code | 60.3 | 23.8 | 2.5× sooner |
One result went the other way. Reading a brand-new prompt, vLLM was faster. Community figures had put the two level.
| Measure | vLLM | SGLang + DFlash2 | Faster |
|---|---|---|---|
| New 19,699-token prompt, tokens per second | 2,154 | 1,819 | vLLM, by 18% |
| Same prompt cached, seconds until the reply starts | 0.5 | 0.2 | SGLang |
The 3D pages in the race are exactly what each engine wrote, unedited, and each one runs in your browser the moment its engine finishes.
Memory settings
The two engines' memory settings measure different things. vLLM's caps everything vLLM uses. SGLang's covers only the model and its cache, and its temporary working memory comes on top (memory creep thread). The Spark's memory is shared with the processor and the operating system, and going past it freezes the whole machine without an error.
| SGLang memory setting | Cache room (est.) | Likely peak (est.) | Assessment |
|---|---|---|---|
| 0.5 | ~1.2M tokens | ~89–104 GB | Headroom to spare. The setting the fastest published Spark recipe uses. |
| 0.6 | ~1.6M tokens | ~102–117 GB | Fits, with less headroom |
| 0.7 | ~2.0M tokens | ~115–130 GB | Can pass the Spark’s memory |
| 0.8 | ~2.4M tokens | ~127–142 GB | Least headroom of the four |
| vLLM | SGLang + DFlash2 | |
|---|---|---|
| Weights | nvidia/Qwen3.8-27B-NVFP4 | RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead |
| Speculative decoding | Off | DFlash2 draft model (z-lab, 2B), 8 tokens per pass |
| Memory setting | 0.7 (gpu-memory-utilization) | 0.5 (mem-fraction-static) |
| Memory in use while writing | 77 GiB of 121 | 68 GiB of 121 |
| Prompt-cache room | 1,270,042 tokens | 557,734 tokens |
| Requests at once | 8 | 6 |
| Test suite | 10 of 10 passed | 10 of 10 passed |
| Known crash triggers | Not needed | 7 of 7 passed |
| Real Claude Code session | Not run in this test | Passed: 2 turns in 26.5 s, both bugs found |
| Engine build | vllm/vllm-openai:qwen38 (3a09141) | lmsysorg/sglang nightly-cu134-20260909 (708f51e) |
Risks, and how to handle them
Speed isn't free on a chip this new to SGLang. The two crash bugs come from community testing (4× Spark write-up); in our run, SGLang passed all 7 of our known-crash-trigger checks.
CLAUDE_CODE_ATTRIBUTION_HEADER=0 where Claude Code is launched.If you want to try it
- Record your current vLLM launch, so you can restart it unchanged.
- Get the SGLang build made for the Spark, the 4-bit (NVFP4) weights and the DFlash2 draft model, about 4 GB.
- Start SGLang on the same port and model name as before, with
--mem-fraction-static 0.5, the Qwen tool-call and reasoning parsers, and metrics on. - Apply the 16-token-minimum patch and add a watchdog that restarts a crashed server.
- For Claude Code, set
CLAUDE_CODE_ATTRIBUTION_HEADER=0and a longAPI_TIMEOUT_MS(SGLang docs). - Before relying on it, test a normal request, a tool call, a request allowing one output token and a cancelled reply.
The limits of these numbers
- Every speed here except ours comes from community testers, and their setups differ.
- Our test is one recorded run per scene with the model's default sampling, so the two engines wrote different text of different lengths. Compare the speeds, not the wording.
- The engines loaded different 4-bit exports of the same model: NVIDIA's for vLLM and RadixArk's for SGLang, whose output layer is full precision. That makes each SGLang pass slightly heavier, not lighter.
- Most tests, ours included, used prompts far shorter than a long coding session builds up, which may shrink the gain in daily use.
On a Spark, the memory ceiling means drafting, not engine tuning, is what moves writing speed: in our test it took a 3D page from 464 seconds to 179. Watch vLLM and SGLang race on one DGX Spark on QAI EdgeBench, with every chart, the setup and the full method.
Sources
- OpenZeka: Comprehensive Qwen3.8-27B study on DGX Sparks (NVIDIA Developer Forums)
- dgx-spark-qwen38: SGLang + 4-bit + DFlash2 recipe
- kubesimplify: Running Qwen3.8-27B on DGX Spark
- 0xBakeer: Qwen3.8-27B 8-bit on a single DGX Spark
- ai-muninn: the DGX Spark bandwidth ceiling
- 4× DGX Spark with SGLang + DFlash2: crash bugs and workarounds (NVIDIA Developer Forums)
- MiaAI-Lab: Qwen3.8-27B SGLang recipe, model and cache sizes
- Inco AI: DFlash 2 release
- z-lab/Qwen3.8-27B-DFlash2 and RadixArk/Qwen3.8-27B-DSpark on Hugging Face
- SGLang docs: Anthropic API and Claude Code setup
- Memory creep on DGX Spark (NVIDIA Developer Forums)