AIOct 4, 2026 · 5 min read · QAI Lab - Boyuan Qian

vLLM vs SGLang on a DGX Spark: Breaking the 14.5 tok/s Wall

SGLang with the DFlash2 draft model wrote Qwen3.8-27B up to 3.3× faster than vLLM on one DGX Spark. Why the Spark has a speed limit, and what beating it costs.

On one NVIDIA DGX Spark, Qwen3.8-27B writes about 12.6 tokens per second under vLLM, and tuning barely moves it, because the hardware caps it near 14.5. With SGLang and the DFlash2 draft model, the same machine wrote 2 to 3.3 times as fast in our recorded test. A 3D page that took vLLM 464 seconds took SGLang 179.

You can watch both engines race, token by token on QAI EdgeBench.

Writing the 3D galaxy page12.6 → 41.2 tok/svLLM → SGLang + DFlash2, same prompt, same Spark
Finish time, 3D page464 s → 179 sSGLang finished 285 s sooner
Editing a file12.6 → 40.0 tok/s3,000 tokens, mean of 2 runs
Test suite10 of 10passed on both engines

Why a Spark tops out near 14.5 tokens per second

To write one token, a language model reads all of its weights from memory. The Spark's memory moves 273 GB per second, and Qwen3.8-27B in 4-bit form is about 18.8 GB, so the ceiling is 273 ÷ 18.8, about 14.5 tokens per second (ai-muninn). The engine doesn't change that arithmetic. Without help, community testers measure every engine on the Spark at 8 to 12 tokens per second, and our own vLLM run came in at 12.6.

In an agentic coding session, most of the wait is the model writing: code, edits, tool calls. So writing speed is the lever.

How speculative decoding gets past it

A small draft model guesses several tokens ahead, eight per pass in DFlash2's case. The main model then checks all of them in a single pass, the same pass it would have spent writing one token, and keeps the ones it would have written itself. Each pass now produces several tokens instead of one. Because the main model still approves every token, the answers are as good as without drafting. Only the speed changes.

The gain depends on how many guesses survive, and that depends on the draft model and on how predictable the text is.

MTP (built in)DSparkDFlash2
Tokens kept per guess, authors’ test4.283.624.80
One Spark, one request18.5 tok/s (vLLM)36.6 tok/s (SGLang)47.9 tok/s (SGLang)
Our test: tokens kept per passnot testednot tested3.8–5.0, by scene
Known problemsnot assessed herenone reported in these sourcesCrashes on very short requests and on cancelled replies
Tokens-kept figures are from DFlash2’s authors (Inco AI), all three measured in one test; speeds from the OpenZeka study. DFlash2 was released on 18 August 2026.

What other Spark owners measured

NVIDIA hasn't published figures for this model on the Spark, so the grey bars below are community results, one request at a time on one Spark. Our own test runs are in blue.

Speed writing a reply, one request at a time, one DGX Spark
Tokens per second. The line marks the limit without speculative decoding.
Our recorded testCommunity results
Limit without drafting · 14.5
llama.cpp4-bit, no drafting
11.6
vLLMour test4-bit, no drafting
12.6
vLLM4-bit + MTP
18.5
Ollama4-bit + MTP, on by default
26.5
vLLM8-bit + DFlash2
31.7
SGLang4-bit + DSpark
36.6
SGLangour test4-bit + DFlash2
41.2
SGLang4-bit + DFlash2
47.9
01020304050
Community setups differ. A separate SGLang + DFlash2 recipe measured 34–40 tokens per second on agentic coding and 18–22 on plain prose: gains depend on how predictable the text is.
Show as table
Setuptok/sSource
llama.cpp, 4-bit, no drafting11.6kubesimplify
vLLM, 4-bit, no drafting12.6QAI EdgeBench, 3D page race
vLLM, 4-bit + MTP18.5OpenZeka study
Ollama, 4-bit + MTP, on by default26.5kubesimplify
vLLM, 8-bit + DFlash231.70xBakeer
SGLang, 4-bit + DSpark36.6OpenZeka study
SGLang, 4-bit + DFlash241.2QAI EdgeBench, 3D page race
SGLang, 4-bit + DFlash247.9OpenZeka study

Without speculative decoding, the engines are about equal, because the memory limit applies to all of them. What separates them is which draft models each one runs well on the Spark.

EngineBest reply speedSetupNotes
SGLang47.9 tok/s4-bit + DFlash2Fastest. 34–40 on agentic coding. DFlash2 has two known crash bugs.
vLLM31.7 tok/s8-bit + DFlash2Measured with 8-bit weights.
Ollama26.5 tok/s4-bit + MTP, on by defaultSimplest to run. Slower at reading prompts: 731 tok/s on 2,000-token prompts.
llama.cpp11.6 tok/s4-bit, no draftingSupports DFlash2, but no Spark measurement was found. 837 tok/s reading 2,000-token prompts.
oMLXn/aApple Silicon onlyDoes not run on the Spark.
Community results as of early October 2026, one request at a time on one Spark. Sources: OpenZeka study, kubesimplify, 0xBakeer.

What we measured

We recorded both engines on 3 October 2026, on the same Spark, with the same model size, the same prompts and the same settings. They ran one after the other, because both servers don't fit in memory at once. vLLM ran without speculative decoding; SGLang ran with DFlash2.

Our test: writing speed by workload
Tokens per second, counted from the first token to the last. Higher is faster.
vLLMSGLang + DFlash2
3D galaxy page, race recording3.3× faster
12.6 tok/s
41.2 tok/s
Editing a file, 3,000 tokens3.2× faster
12.6 tok/s
40.0 tok/s
Writing new code, 3,000 tokens2.2× faster
12.7 tok/s
28.0 tok/s
Python tutorial, 512 tokens2.0× faster
12.7 tok/s
24.9 tok/s
Show as table
vLLM, tok/sSGLang + DFlash2, tok/sDifferenceSource
3D galaxy page, race recording12.641.23.3× fasterRace recording
Editing a file, 3,000 tokens12.640.03.2× fasterBenchmark run, mean of 2
Writing new code, 3,000 tokens12.728.02.2× fasterBenchmark run, mean of 2
Python tutorial, 512 tokens12.724.92.0× fasterDecode check

Editing a file landed at the top of the community's 34 to 40 range, and writing new code fell below it. That fits how drafting works: an edit mostly repeats text the model has just read, so more of the draft survives. The finish times show what that means for waiting.

Our test: time to finish each race scene
Seconds from sending the request to the last token. Lower is faster.
vLLMSGLang + DFlash2
3D galaxy page2.6× sooner
463.9 s
178.8 s
Editing a file3.1× sooner
60.7 s
19.7 s
Writing new code2.5× sooner
60.3 s
23.8 s
Show as table
vLLM, sSGLang + DFlash2, sDifference
3D galaxy page463.9178.82.6× sooner
Editing a file60.719.73.1× sooner
Writing new code60.323.82.5× sooner

One result went the other way. Reading a brand-new prompt, vLLM was faster. Community figures had put the two level.

Our test: reading the prompt
MeasurevLLMSGLang + DFlash2Faster
New 19,699-token prompt, tokens per second2,1541,819vLLM, by 18%
Same prompt cached, seconds until the reply starts0.50.2SGLang

The 3D pages in the race are exactly what each engine wrote, unedited, and each one runs in your browser the moment its engine finishes.

Memory settings

The two engines' memory settings measure different things. vLLM's caps everything vLLM uses. SGLang's covers only the model and its cache, and its temporary working memory comes on top (memory creep thread). The Spark's memory is shared with the processor and the operating system, and going past it freezes the whole machine without an error.

SGLang memory settingCache room (est.)Likely peak (est.)Assessment
0.5~1.2M tokens~89–104 GBHeadroom to spare. The setting the fastest published Spark recipe uses.
0.6~1.6M tokens~102–117 GBFits, with less headroom
0.7~2.0M tokens~115–130 GBCan pass the Spark’s memory
0.8~2.4M tokens~127–142 GBLeast headroom of the four
Estimates use published sizes: 19.7 GB for the 4-bit model, 4 GB for the DFlash2 draft model and 32.8 KB of cache per token, plus 25–40 GB of temporary memory one tester measured. In our test at 0.5, SGLang reported 557,734 tokens of cache room, about half the estimate, and used 68 GiB while writing.
Our test: setup and reliability
vLLMSGLang + DFlash2
Weightsnvidia/Qwen3.8-27B-NVFP4RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead
Speculative decodingOffDFlash2 draft model (z-lab, 2B), 8 tokens per pass
Memory setting0.7 (gpu-memory-utilization)0.5 (mem-fraction-static)
Memory in use while writing77 GiB of 12168 GiB of 121
Prompt-cache room1,270,042 tokens557,734 tokens
Requests at once86
Test suite10 of 10 passed10 of 10 passed
Known crash triggersNot needed7 of 7 passed
Real Claude Code sessionNot run in this testPassed: 2 turns in 26.5 s, both bugs found
Engine buildvllm/vllm-openai:qwen38 (3a09141)lmsysorg/sglang nightly-cu134-20260909 (708f51e)

Risks, and how to handle them

Speed isn't free on a chip this new to SGLang. The two crash bugs come from community testing (4× Spark write-up); in our run, SGLang passed all 7 of our known-crash-trigger checks.

CrashVery short requests stop the server
A request that allows 8 or fewer output tokens stalls SGLang with DFlash2 and kills it.
Fix: a community patch that raises the minimum to 16 tokens.
CrashCancelling a reply can stop the server
Stopping a streaming reply, for example by pressing Esc in Claude Code, can crash SGLang with DFlash2. There is no upstream fix yet.
Mitigation: a watchdog that restarts a crashed server. Restarts took about 6–8 minutes on one 4-Spark cluster. If it shows up in testing, switch the draft model to DSpark.
FreezeMemory set too high locks the Spark
Past the Spark’s memory, the whole machine hard-locks without a log, or the memory killer stops SGLang.
Mitigation: start at memory setting 0.5 and watch peak memory through a long session.
SlowerLosing the prompt cache
By default, Claude Code adds a per-request marker near the start of its prompt, which stops SGLang reusing the cached prompt, so every turn rereads the whole prompt.
Fix: set CLAUDE_CODE_ATTRIBUTION_HEADER=0 where Claude Code is launched.
SetupTool calls arrive as plain text
SGLang needs the Qwen tool-call and reasoning parsers set. With the wrong parser, Claude Code’s tools stop working, which looks like a weaker model.
Fix: match the parsers your current server uses, and include a tool call in your tests.
MaturityThe Spark’s chip is still new to SGLang
SGLang’s support for the Spark’s chip is still incomplete, some 4-bit setups have crashed, and the working recipes are maintained by the community.

If you want to try it

  1. Record your current vLLM launch, so you can restart it unchanged.
  2. Get the SGLang build made for the Spark, the 4-bit (NVFP4) weights and the DFlash2 draft model, about 4 GB.
  3. Start SGLang on the same port and model name as before, with --mem-fraction-static 0.5, the Qwen tool-call and reasoning parsers, and metrics on.
  4. Apply the 16-token-minimum patch and add a watchdog that restarts a crashed server.
  5. For Claude Code, set CLAUDE_CODE_ATTRIBUTION_HEADER=0 and a long API_TIMEOUT_MS (SGLang docs).
  6. Before relying on it, test a normal request, a tool call, a request allowing one output token and a cancelled reply.

The limits of these numbers

  • Every speed here except ours comes from community testers, and their setups differ.
  • Our test is one recorded run per scene with the model's default sampling, so the two engines wrote different text of different lengths. Compare the speeds, not the wording.
  • The engines loaded different 4-bit exports of the same model: NVIDIA's for vLLM and RadixArk's for SGLang, whose output layer is full precision. That makes each SGLang pass slightly heavier, not lighter.
  • Most tests, ours included, used prompts far shorter than a long coding session builds up, which may shrink the gain in daily use.

On a Spark, the memory ceiling means drafting, not engine tuning, is what moves writing speed: in our test it took a 3D page from 464 seconds to 179. Watch vLLM and SGLang race on one DGX Spark on QAI EdgeBench, with every chart, the setup and the full method.

Sources