系列:vLLM & SGLang Serving Notes

(6) Supported Models & Selection Guide — Who to Pick for What Scenario

1. Model Support Matrix

Conclusion first: both cover mainstream models well; the difference is “how fast a new model is usable” and “how deep the tuning goes”.

ModelScale / structurevLLMSGLangNote
DeepSeek-V4 / V4 ProMoE + MLA + DSAFirst-class, native DSpark specFirst-class, official cookbook benchmark8×B300 ≈ 250 tok/s, DSpark 12–42% above MTP
GLM-5.2MoE + DSAProduction plan: NVFP4+MTP+P/DDeep-tuning benchmarkBlackwell 500+ tok/s/user; PCP prefill 20.1k→27.3k
Qwen3 / Qwen3.5 / Qwen3-VLdense + MoE + VLMFirst-class (incl Transformers backend M-RoPE)First-classWatch GPTQ + spec combo trap (#48816)
Step 3.7 Flash196B MoE / 11B active / 256KPrebuilt image vllm/vllm-openai:stepfun37Dev image lmsysorg/sglang:dev-step-3.7-flashFP8 / BF16 / NVFP4 + MTP / EAGLE
Step-3.5-Flash196B MoE + 3:1 sliding windowOfficial recipe deploy guideSupportedInt4 weights not yet in vLLM
Step3-VL-10B10B edge VLMnightly ≥ 0.14.0rc2latest main + cookbookSingle RTX 4090, AIME2025 94.43%
Tencent Hy3295B MoEDay-oneDay-oneDomestic model day-one case
Llama / Mistral / GemmadenseFullFullGemma4: NVFP4 needs 0.25.1+
Cosmos3 Edge videomultimodalSupported (#49190 fix)—vLLM wider multimodal
New arch just out on HFanyDay-0 full speed (Transformers parity)Wait for nativevLLM clear edge
The single most important difference: vLLM's Transformers backend parity means as long as HF has an implementation, a new model is served at full speed the same day — no waiting for the framework to write a native kernel. For teams chasing new models, this one point often decides the selection.

2. Hardware Support Matrix

HardwarevLLMSGLangNote
NVIDIA Hopper (H100/H200)✅ mature✅ matureFP8 native
NVIDIA Blackwell (B200/B300/GB200)✅ FA4 + SM100 FP8 KV✅ NVFP4 deep tuningSGLang more aggressive on NVFP4
Consumer Blackwell (SM120)✅✅ #30272 DeepSeek-V4 mxfp4 MoE + TP2Single / dual-card path
AMD ROCm✅ AITER v0.1.16.post5✅ MXFP4 (#28291)vLLM wider
Intel XPU✅ DeepSeek-V4 fuse_index_q SYCLLimitedvLLM exclusive edge
TPU / CPU✅LimitedvLLM leads breadth

Conclusion: on hardware breadth vLLM leads clearly; but on the most frontier combo, Blackwell + NVFP4, SGLang tunes deeper.

3. Selection Decision Tree

vLLM or SGLang? Do requests share a long prefix? (system / multi-turn / eval) Yes, high share No / unsure Prefer SGLang Need to run a just-released model? · Agent / multi-turn chat · Batch eval / tree-of-thought · Strict structured output · Large EP · NVFP4 extreme tuning Yes No Choose vLLM (Day-0 full speed) Non-NVIDIA hardware? Yes No Choose vLLM (ROCm/XPU/TPU) Either works Practical advice: it is not a single-choice question · Both expose OpenAI-compatible APIs → low switching cost; worth benchmarking each with your real traffic before deciding · Large-scale setups can co-deploy: Agent main path on SGLang (prefix reuse), long-tail / new models on vLLM (breadth) · Benchmarking must look at TTFT / P99 TPOT / throughput; total throughput alone leads to wrong picks

4. One-Line Selection Mantra

Heavy shared prefix → SGLang; need model / hardware breadth → vLLM.

Spelled out a bit:

Signals leaning SGLang

Signals leaning vLLM

5. Deployment Notes for Typical Models

5.1 Step 3.7 Flash (196B MoE / 11B active / 256K)

5.2 Step-3.5-Flash: why it’s both fast and lean

This model is a good example of “three tricks stacked”:

  1. 3:1 hybrid sliding-window attention: most layers use the sliding window, dropping the main cost from O(n²) to O(n·w); a few global layers handle long-range information flow;
  2. MTP built-in draft head: ~4 tokens per step, i.e. built-in speculative decoding;
  3. Sparse MoE: 196B total params, only 11B active per token — memory pays total params, compute pays active.
Deployment note: on vLLM, MoE / MTP related optimizations are NOT all on by default — they need explicit config per the official recipe; Int4 weights are not yet supported in vLLM.

5.3 GLM-5.2: currently the deepest-tuned combo

The production plan stacks three pieces: NVFP4 quant + MTP spec + P/D disagg.

5.4 Step3-VL-10B: the edge-multimodal landing point

6. Deployment Practice Checklist

Before launch

Suggested tuning order

  1. Set --max-model-len to the real need (directly decides the KV budget)
  2. Enable prefix cache / RadixAttention (almost no side effects)
  3. Tune --max-num-seqs and gpu-memory-utilization for the throughput knee
  4. Then enable speculative decoding (MTP / EAGLE), watch acceptance > 60%
  5. Only then consider quant and PD disaggregation (big gain but big complexity)

When you don’t need PD disaggregation

7. Series Summary

Six posts in, the main line looks like this:

#TopicOne line
1Why an engine is neededNaive inference wastes on “memory holes, slot idle, GPU waiting”
2vLLM internalsStarted with paged KV + continuous batching; kills sync via MRv2, stands on breadth
3SGLang internalsStarts from “an LLM app is a structured program”; stands on prefix reuse + frontier throughput
4The shared frontierFour battle lines — sync stall, spec decode, PD disagg, low-bit quant — stack on each other
5Version evolutionJuly 2026, both nearly synchronized an architecture turnover; a big release is always followed by fixes
6Models & selectionHeavy prefix → SGLang, breadth → vLLM; best to benchmark each with real traffic
Last observation: the two engines' competition is purely good for users. One ships MRv2, the other soon has Spec V2; one goes NVFP4, the other immediately follows with NVFP4_AWQ. Both expose OpenAI-compatible APIs, so switching cost is low — don't treat selection as a one-time lifelong decision; periodically re-benchmark with your own real traffic.

Related reading: the algorithm behind speculative decoding is in the Speculative Decoding Notes; model-side MoE, attention, and multimodal structures are in the Multimodal Decoding Notes.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。