vLLM vs SGLang: Qwen3.8-27B on DGX Spark

SGLang with the DFlash2 draft model wrote 2.0 to 3.3 times as fast as vLLM on the same machine. vLLM read a brand-new prompt 18% faster.

Model
Qwen3.8-27B · 4-bit (NVFP4)
Hardware
NVIDIA DGX Spark · GB10 · 121 GiB unified memory · 273 GB/s
Recorded
Speed

Prompt: write a full-screen three.js spiral galaxy page: 20,000+ glowing particles, camera orbit, mouse parallax, starfield · up to 8,192 tokens · thinking off

vLLM

Standard: 4-bit weights, no speculative decoding

Elapsed
–
Tokens written
–
Writing speed
–
Loading the recording…
The page runs here as soon as this engine finishes writing it.

SGLang + DFlash2

Accelerated: 4-bit weights, DFlash2 draft model, 8 tokens per pass

Elapsed
–
Tokens written
–
Writing speed
–
Loading the recording…
The page runs here as soon as this engine finishes writing it.

Finish line: SGLang + DFlash2 finished 2.6× sooner, 285.2 seconds ahead.

vLLM 463.9 s
SGLang + DFlash2 178.8 s
0 s100 s200 s300 s400 s500 s

Each engine got the identical request; these are its real token streams, replayed at the speed it wrote them. Recorded one after the other on the same Spark, because both servers do not fit in memory at once.

Writing the 3D page
vLLM
12.6 tok/s
SGLang
41.2 tok/s
SGLang 3.3× faster
Finish time, 3D page
vLLM
464 s
SGLang
179 s
SGLang done 285 s sooner
Editing a file
vLLM
12.6 tok/s
SGLang
40.0 tok/s
SGLang 3.2× faster
Reading a new prompt
vLLM
2,154 tok/s
SGLang
1,819 tok/s
vLLM 18% faster

Inside the race

The scene selected above (3D galaxy page), second by second.

Tokens written over time

A steeper line means faster writing.

vLLMSGLang + DFlash2
Loading the recording…

Writing speed over time

Tokens per second, averaged over a rolling 5-second window.

vLLMSGLang + DFlash2
Loading the recording…

Output speed

Tokens written per second, counted from the first token to the last, thinking included. vLLM stays close to the Spark's memory-bandwidth limit of about 14.5 tokens per second. SGLang gets past it by checking several drafted tokens in each pass.

Writing speed by workload

Tokens per second. Higher is faster.

vLLMSGLang + DFlash2
3D galaxy page, race recording3.3× faster
Editing a file, 3,000 tokens3.2× faster
Writing new code, 3,000 tokens2.2× faster
Python tutorial, 512 tokens2.0× faster
Show as table
WorkloadvLLM, tok/sSGLang + DFlash2, tok/sSpeed-upSource
3D galaxy page, race recording12.641.23.3×Race recording
Editing a file, 3,000 tokens12.640.03.2×Benchmark run, mean of 2
Writing new code, 3,000 tokens12.728.02.2×Benchmark run, mean of 2
Python tutorial, 512 tokens12.724.92.0×Decode check

Input speed

How fast each engine reads the prompt before it starts writing. vLLM reads a brand-new prompt faster. Both reuse a cached prompt, and SGLang returns from its cache sooner.

Reading a new prompt

A new 19,699-token prompt, in tokens per second. Higher is faster.

vLLMSGLang + DFlash2
New promptvLLM 18% faster

Reading the same prompt again

Seconds until the reply starts, with the prompt cached. Lower is faster.

vLLMSGLang + DFlash2
Same prompt, cached

First token in each race scene

Seconds from sending the request to the first token, with the prompt cached on both. Lower is faster.

vLLMSGLang + DFlash2
3D galaxy page
Editing a file
Writing new code

Where the speed comes from

In every pass the DFlash2 draft model proposes 8 tokens, and the main model keeps the ones it would have written itself. Each dot is SGLang's average over a 5-second stretch of the recording. vLLM writes one token per pass.

Tokens kept per pass

Higher means more of the draft was kept.

3D galaxy page37 stretches, 3.55–5.97
Editing a file4 stretches, 3.83–5.65
Writing new code5 stretches, 3.42–4.38
012345678vLLM: 1 per pass

Setup and reliability

MeasurevLLMSGLang + DFlash2
Weightsnvidia/Qwen3.8-27B-NVFP4RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead
Speculative decodingOffDFlash2 draft model (z-lab, 2B), 8 tokens per pass
Memory setting0.7 (gpu-memory-utilization)0.5 (mem-fraction-static)
Memory in use while writing77 GiB of 12168 GiB of 121
Prompt-cache room1,270,042 tokens557,734 tokens
Requests at once86
Test suite10 of 10 passed10 of 10 passed
Known crash triggersNot needed7 of 7 passed
Real Claude Code sessionNot run in this testPassed: 2 turns in 26.5 s, both bugs found
Engine buildvllm/vllm-openai:qwen38 (3a09141)lmsysorg/sglang nightly-cu134-20260909 (708f51e)

How this was measured

  • Same DGX Spark (GB10, 121 GiB unified memory, 273 GB/s), same model size (Qwen3.8-27B, 4-bit), same API: Anthropic /v1/messages, the way Claude Code calls it.
  • Each scene used the identical prompt, token limit and Claude Code effort, with the model's default sampling. Each prompt was sent once beforehand, so both engines started with it cached.
  • The coding scenes went through Anthropic /v1/messages with Claude Code effort "medium". The 3D page went through /v1/chat/completions with thinking switched off on both engines: in a first recording with thinking on, vLLM spent its whole 6,144-token budget reasoning and wrote no page.
  • The engines ran one after the other on 2026-10-03, because both servers do not fit in memory together. Nothing else was running on the Spark.
  • vLLM ran in its standard configuration, without speculative decoding. SGLang ran with the DFlash2 draft model at memory setting 0.5.
  • The engines load different 4-bit exports of the same model: NVIDIA's for vLLM, and RadixArk's for SGLang, whose output layer is full precision. That output layer makes each SGLang pass slightly heavier, not lighter.
  • Speeds count every token written, thinking included, from the first token to the last. Token totals come from each server's own usage report.
  • Sampling is random, so the two engines wrote different text of different lengths. Compare the speeds, not the wording.
  • The 3D pages are exactly what each engine wrote, unedited, running in a sandboxed frame.