1. Model Support Matrix
Conclusion first: both cover mainstream models well; the difference is “how fast a new model is usable” and “how deep the tuning goes”.
| Model | Scale / structure | vLLM | SGLang | Note |
|---|---|---|---|---|
| DeepSeek-V4 / V4 Pro | MoE + MLA + DSA | First-class, native DSpark spec | First-class, official cookbook benchmark | 8×B300 ≈ 250 tok/s, DSpark 12–42% above MTP |
| GLM-5.2 | MoE + DSA | Production plan: NVFP4+MTP+P/D | Deep-tuning benchmark | Blackwell 500+ tok/s/user; PCP prefill 20.1k→27.3k |
| Qwen3 / Qwen3.5 / Qwen3-VL | dense + MoE + VLM | First-class (incl Transformers backend M-RoPE) | First-class | Watch GPTQ + spec combo trap (#48816) |
| Step 3.7 Flash | 196B MoE / 11B active / 256K | Prebuilt image vllm/vllm-openai:stepfun37 | Dev image lmsysorg/sglang:dev-step-3.7-flash | FP8 / BF16 / NVFP4 + MTP / EAGLE |
| Step-3.5-Flash | 196B MoE + 3:1 sliding window | Official recipe deploy guide | Supported | Int4 weights not yet in vLLM |
| Step3-VL-10B | 10B edge VLM | nightly ≥ 0.14.0rc2 | latest main + cookbook | Single RTX 4090, AIME2025 94.43% |
| Tencent Hy3 | 295B MoE | Day-one | Day-one | Domestic model day-one case |
| Llama / Mistral / Gemma | dense | Full | Full | Gemma4: NVFP4 needs 0.25.1+ |
| Cosmos3 Edge video | multimodal | Supported (#49190 fix) | — | vLLM wider multimodal |
| New arch just out on HF | any | Day-0 full speed (Transformers parity) | Wait for native | vLLM clear edge |
Transformers backend parity means as long as HF has an implementation, a new model is served at full speed the same day — no waiting for the framework to write a native kernel. For teams chasing new models, this one point often decides the selection.
2. Hardware Support Matrix
| Hardware | vLLM | SGLang | Note |
|---|---|---|---|
| NVIDIA Hopper (H100/H200) | ✅ mature | ✅ mature | FP8 native |
| NVIDIA Blackwell (B200/B300/GB200) | ✅ FA4 + SM100 FP8 KV | ✅ NVFP4 deep tuning | SGLang more aggressive on NVFP4 |
| Consumer Blackwell (SM120) | ✅ | ✅ #30272 DeepSeek-V4 mxfp4 MoE + TP2 | Single / dual-card path |
| AMD ROCm | ✅ AITER v0.1.16.post5 | ✅ MXFP4 (#28291) | vLLM wider |
| Intel XPU | ✅ DeepSeek-V4 fuse_index_q SYCL | Limited | vLLM exclusive edge |
| TPU / CPU | ✅ | Limited | vLLM leads breadth |
Conclusion: on hardware breadth vLLM leads clearly; but on the most frontier combo, Blackwell + NVFP4, SGLang tunes deeper.
3. Selection Decision Tree
4. One-Line Selection Mantra
Spelled out a bit:
Signals leaning SGLang
- Main scenario is Agent, multi-turn dialogue, ReAct loops (system prompt + tool defs resent every step)
- Batch eval, tree-of-thought, self-consistency sampling (one prefix, many forks)
- Strict JSON Schema / regex constrained output, and high share
- Running DeepSeek / GLM-style MoE + MLA/DSA models, wanting NVFP4 extreme throughput
- Needing large-scale expert-parallel (EP) deployment
Signals leaning vLLM
- Diverse model types, frequently needing just-released new architectures
- Hardware not pure NVIDIA (ROCm / Intel XPU / TPU)
- Diverse quant-format needs (AWQ / GPTQ / FP8 / NVFP4 / MXFP4 mixed)
- Want to plug into existing ecosystem: KV Connector, Rust router, Dynamo, llm-d
- Team values “stable version, complete docs, pitfalls already stepped on”
5. Deployment Notes for Typical Models
5.1 Step 3.7 Flash (196B MoE / 11B active / 256K)
- vLLM: the official prebuilt image
vllm/vllm-openai:stepfun37is the most stable, supports MTP spec + NVFP4 4-card deployment - SGLang:
lmsysorg/sglang:dev-step-3.7-flash+ EAGLE - Both frameworks’ adaptation is “first-class mature”, but deep tuning and public-benchmark maturity still lag GLM / DeepSeek
- Model side requires
transformers ≥ 5.0(custom modeling, viatrust_remote_code)
5.2 Step-3.5-Flash: why it’s both fast and lean
This model is a good example of “three tricks stacked”:
- 3:1 hybrid sliding-window attention: most layers use the sliding window, dropping the main cost from O(n²) to O(n·w); a few global layers handle long-range information flow;
- MTP built-in draft head: ~4 tokens per step, i.e. built-in speculative decoding;
- Sparse MoE: 196B total params, only 11B active per token — memory pays total params, compute pays active.
5.3 GLM-5.2: currently the deepest-tuned combo
The production plan stacks three pieces: NVFP4 quant + MTP spec + P/D disagg.
IndexerCachelifts MTP acceptance- PCP (context parallel) lifts prefill throughput from 20.1k to 27.3k
- SGLang side: IndexShare MTP (draft reuses top-k, long-context cost cut 1.9×) + TopK-V2 (Lightning-TopK, optimizes 80k-level inputs)
- result: Blackwell 500+ tok/s/user
- recommend locking SGLang v0.5.15.post1 or higher
5.4 Step3-VL-10B: the edge-multimodal landing point
- 10B VLM (PE-lang 1.8B + Qwen3-8B), runnable on a single RTX 4090 (BF16 / FP8)
- AIME2025 94.43%
- needs vLLM nightly ≥ 0.14.0rc2 or SGLang latest main (official cookbook)
- depends on nightly / main, production must lock the version
- currently the most worth-following landing point for domestic edge multimodal inference
6. Deployment Practice Checklist
Before launch
- Benchmark with real traffic distribution, not fixed-length synthetic requests
- Measure all three metrics: TTFT, P99 TPOT, total throughput
- Quant models run an output-correctness regression (NVFP4 garbage is silent)
- Confirm the prefix-sharing rate — decides whether RadixAttention pays off for you
- Lock the release tag, don’t follow main
Suggested tuning order
- Set
--max-model-lento the real need (directly decides the KV budget) - Enable prefix cache / RadixAttention (almost no side effects)
- Tune
--max-num-seqsandgpu-memory-utilizationfor the throughput knee - Then enable speculative decoding (MTP / EAGLE), watch acceptance > 60%
- Only then consider quant and PD disaggregation (big gain but big complexity)
When you don’t need PD disaggregation
- Single 8-GPU box or smaller, small/medium model → colocated deployment is simpler
- Prompts generally short → prefill doesn’t interfere anyway
- No NVLink / RDMA fast interconnect → KV transfer eats the gains
7. Series Summary
Six posts in, the main line looks like this:
| # | Topic | One line |
|---|---|---|
| 1 | Why an engine is needed | Naive inference wastes on “memory holes, slot idle, GPU waiting” |
| 2 | vLLM internals | Started with paged KV + continuous batching; kills sync via MRv2, stands on breadth |
| 3 | SGLang internals | Starts from “an LLM app is a structured program”; stands on prefix reuse + frontier throughput |
| 4 | The shared frontier | Four battle lines — sync stall, spec decode, PD disagg, low-bit quant — stack on each other |
| 5 | Version evolution | July 2026, both nearly synchronized an architecture turnover; a big release is always followed by fixes |
| 6 | Models & selection | Heavy prefix → SGLang, breadth → vLLM; best to benchmark each with real traffic |
Related reading: the algorithm behind speculative decoding is in the Speculative Decoding Notes; model-side MoE, attention, and multimodal structures are in the Multimodal Decoding Notes.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。