vLLM vs SGLang: Qwen3.8-27B on DGX Spark
SGLang with the DFlash2 draft model wrote 2.0 to 3.3 times as fast as vLLM on the same machine. vLLM read a brand-new prompt 18% faster.
- Model
- Qwen3.8-27B · 4-bit (NVFP4)
- Hardware
- NVIDIA DGX Spark · GB10 · 121 GiB unified memory · 273 GB/s
- Recorded
Prompt: write a full-screen three.js spiral galaxy page: 20,000+ glowing particles, camera orbit, mouse parallax, starfield · up to 8,192 tokens · thinking off
vLLM
Finished in 463.9 sStandard: 4-bit weights, no speculative decoding
- Elapsed
- –
- Tokens written
- –
- Writing speed
- –
Loading the recording…THINKING ANSWER
SGLang + DFlash2
Finished in 178.8 sAccelerated: 4-bit weights, DFlash2 draft model, 8 tokens per pass
- Elapsed
- –
- Tokens written
- –
- Writing speed
- –
Loading the recording…THINKING ANSWER
Finish line: SGLang + DFlash2 finished 2.6× sooner, 285.2 seconds ahead.
Each engine got the identical request; these are its real token streams, replayed at the speed it wrote them. Recorded one after the other on the same Spark, because both servers do not fit in memory at once.
- vLLM
- 12.6 tok/s
- SGLang
- 41.2 tok/s
- vLLM
- 464 s
- SGLang
- 179 s
- vLLM
- 12.6 tok/s
- SGLang
- 40.0 tok/s
- vLLM
- 2,154 tok/s
- SGLang
- 1,819 tok/s
Inside the race
The scene selected above (3D galaxy page), second by second.
Tokens written over time
A steeper line means faster writing.
Writing speed over time
Tokens per second, averaged over a rolling 5-second window.
Output speed
Tokens written per second, counted from the first token to the last, thinking included. vLLM stays close to the Spark's memory-bandwidth limit of about 14.5 tokens per second. SGLang gets past it by checking several drafted tokens in each pass.
Writing speed by workload
Tokens per second. Higher is faster.
Show as table
| Workload | vLLM, tok/s | SGLang + DFlash2, tok/s | Speed-up | Source |
|---|---|---|---|---|
| 3D galaxy page, race recording | 12.6 | 41.2 | 3.3× | Race recording |
| Editing a file, 3,000 tokens | 12.6 | 40.0 | 3.2× | Benchmark run, mean of 2 |
| Writing new code, 3,000 tokens | 12.7 | 28.0 | 2.2× | Benchmark run, mean of 2 |
| Python tutorial, 512 tokens | 12.7 | 24.9 | 2.0× | Decode check |
Input speed
How fast each engine reads the prompt before it starts writing. vLLM reads a brand-new prompt faster. Both reuse a cached prompt, and SGLang returns from its cache sooner.
Reading a new prompt
A new 19,699-token prompt, in tokens per second. Higher is faster.
Reading the same prompt again
Seconds until the reply starts, with the prompt cached. Lower is faster.
First token in each race scene
Seconds from sending the request to the first token, with the prompt cached on both. Lower is faster.
Where the speed comes from
In every pass the DFlash2 draft model proposes 8 tokens, and the main model keeps the ones it would have written itself. Each dot is SGLang's average over a 5-second stretch of the recording. vLLM writes one token per pass.
Tokens kept per pass
Higher means more of the draft was kept.
Setup and reliability
| Measure | vLLM | SGLang + DFlash2 |
|---|---|---|
| Weights | nvidia/Qwen3.8-27B-NVFP4 | RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead |
| Speculative decoding | Off | DFlash2 draft model (z-lab, 2B), 8 tokens per pass |
| Memory setting | 0.7 (gpu-memory-utilization) | 0.5 (mem-fraction-static) |
| Memory in use while writing | 77 GiB of 121 | 68 GiB of 121 |
| Prompt-cache room | 1,270,042 tokens | 557,734 tokens |
| Requests at once | 8 | 6 |
| Test suite | 10 of 10 passed | 10 of 10 passed |
| Known crash triggers | Not needed | 7 of 7 passed |
| Real Claude Code session | Not run in this test | Passed: 2 turns in 26.5 s, both bugs found |
| Engine build | vllm/vllm-openai:qwen38 (3a09141) | lmsysorg/sglang nightly-cu134-20260909 (708f51e) |
How this was measured
- Same DGX Spark (GB10, 121 GiB unified memory, 273 GB/s), same model size (Qwen3.8-27B, 4-bit), same API: Anthropic /v1/messages, the way Claude Code calls it.
- Each scene used the identical prompt, token limit and Claude Code effort, with the model's default sampling. Each prompt was sent once beforehand, so both engines started with it cached.
- The coding scenes went through Anthropic /v1/messages with Claude Code effort "medium". The 3D page went through /v1/chat/completions with thinking switched off on both engines: in a first recording with thinking on, vLLM spent its whole 6,144-token budget reasoning and wrote no page.
- The engines ran one after the other on 2026-10-03, because both servers do not fit in memory together. Nothing else was running on the Spark.
- vLLM ran in its standard configuration, without speculative decoding. SGLang ran with the DFlash2 draft model at memory setting 0.5.
- The engines load different 4-bit exports of the same model: NVIDIA's for vLLM, and RadixArk's for SGLang, whose output layer is full precision. That output layer makes each SGLang pass slightly heavier, not lighter.
- Speeds count every token written, thinking included, from the first token to the last. Token totals come from each server's own usage report.
- Sampling is random, so the two engines wrote different text of different lengths. Compare the speeds, not the wording.
- The 3D pages are exactly what each engine wrote, unedited, running in a sandboxed frame.