Daily AI Hotspot Notes · All 27 Episodes
Series · 27 episodes (ongoing). Front-line daily progress on vLLM / SGLang upstream, domestic-model enablement, speculative decoding and inference optimization, framed as "same-day same-frame / who is falling behind"; each post has a Chinese counterpart.
- Daily AI Hotspot · 2026-09-27: No New Release on Sunday — SGLang main Closes Cache/Routing Base, kv-hints Envelope Enters Transport; DriveVLM Puts VLM Slow-Thinking Into the Autonomous-Driving Loop 2026-09-27T00:00:00.000Z
- Daily AI Hotspot · 2026-09-26: vLLM Holds v0.30.0 and Cuts v0.30.1rc0 (AMD MI355 NVFP4); Step-5-Preview-BF16 Weights Quietly Appear on HF as the Most Notable Domestic Inference Progress; MAE Founds Self-Supervised Visual Front-End 2026-09-26T00:00:00.000Z
- Daily AI Hotspot · 2026-09-25: vLLM v0.30.0 Annual Major Release (762 commits, Fast Start Persistent Weight Cache + HiSparse Host KV Spillover, CVE-2026-93436 Fix Landed); DINOv3 7B Frozen Visual Backbone Beats Weak-Supervision SOTA for the First Time 2026-09-25T00:00:00.000Z
- Daily AI Hotspot · 2026-09-21: Step 5 Preview Released (600B/27B Ultra-Sparse MoE, 1M Context, Open Weights 10-15); SGLang v0.5.20 Sampling Masks Overlap Lets RL Replay Rollouts, Qwen3-8B Decode +52% 2026-09-21T00:00:00.000Z
- Daily AI Hotspot · 2026-09-20: NVIDIA GR00T N2 Upgrades the Robot Brain From VLA to a World Action Model (2× Success on New Environments); SGLang v0.5.20 HRRN Scheduling Cuts TTFT 69%, NVIDIA Dynamo KV-Aware Routing Cuts TTFT −50% 2026-09-20T00:00:00.000Z
- Daily AI Hotspot · 2026-09-19: SGLang v0.5.20 Ships (713 PRs), #37709 DSpark-under-PD Makes Draft KV Reusable Across Nodes; Decision Transformer Anchors VLA's 'Action as Token' Lineage 2026-09-19T00:00:00.000Z
- Daily AI Hotspot · 2026-09-17: vLLM #56935 Makes DSv4.1 Fused Mega Attention + NVFP4 Compressed KV SM100 Default; π*0.6/RECAP Shifts Robot Deployment from 'Model Strength' to 'Real Deploy Data' 2026-09-17T00:00:00.000Z
- Daily AI Hotspot · 2026-09-16: StepAudio 3 Tops Artificial Analysis + Step-3.7-Flash Deploy Path Matures, SigLIP 2 Becomes the Default VLA Vision Tower 2026-09-16T00:00:00.000Z
- Daily AI Hotspot · 2026-09-15: vLLM v0.29.1rc0 Ships Speculative-Decode Dual-Key Watermark + EPD Three-Segment Split, SGLang main Adds Diffusion & XPU DFlash 2026-09-15T00:00:00.000Z
- Daily AI Hotspot · 2026-09-13: vLLM HiSparse Trio Turns Host Memory into a GPU VRAM Extension Layer, SGLang graph-pool Makes CUDA Graph VRAM a Borrowable Pool 2026-09-13T00:00:00.000Z
- Daily AI Hotspot · 2026-09-12: SGLang Lands Day-0 DeepSeek-V4.1 Support (Bounded Recompute for Less Cache — 8×H200 Prefill 1.56×, pass@1 Lossless), vLLM Holds at v0.29.0 2026-09-12T00:00:00.000Z
- Daily AI Hotspot · 2026-09-10: vLLM Ships v0.29.0 (MRV2 Default + Mamba Prefix Cache −9~25% TTFT), SGLang Holds at v0.5.19 2026-09-10T00:00:00.000Z
- Daily AI Hotspot · 2026-09-07: SGLang Ships v0.5.19 (786 PRs — W4A8 MoE +12% / DeepEP v2 Decode into CUDA Graph), vLLM Stuck in 0.29.0 Candidate with Two Security Advisories on Stable 2026-09-07T00:00:00.000Z
- Daily AI Hotspot · 2026-09-03: Agentic Workloads Become the Main Battleground — The Win Condition Shifts From Fixed 8k Throughput to Whether a Multi-Turn Session Still Hits Its History KV 2026-09-03T00:00:00.000Z
- Daily AI Hotspot · 2026-09-02: vLLM's Adaptive Speculative Budget Cuts DSpark TTFT 55–65% at Zero Model Cost; RT-X Confirms Robots Obey a Data Scaling Law 2026-09-02T00:00:00.000Z
- Daily AI Hotspot · 2026-09-01: SGLang Turns NVFP4 + FP8 lm_head Into Endless Repetition — And You Cannot Install the Fix on v0.5.18 2026-09-01T00:00:00.000Z
- Daily AI Hotspot · 2026-08-29: SGLang Cuts Cold Start to 35.6s (2.38×), and vLLM v0.28.0's Three Breaking Changes 2026-08-29T00:00:00.000Z
- Daily AI Hotspot · 2026-08-28: Three Default Values Changed on main After v0.28.0, While SGLang Rebuilds Its Prefill Foundation 2026-08-28T00:00:00.000Z
- Daily AI Hotspot · 2026-08-27: vLLM v0.28.0 Ships (584 Commits) — Full-Stack Kimi-K3 Performance, End-to-End DeepSeek V4 Sparse MLA, DFlash2 Reaches Stable 2026-08-27T00:00:00.000Z
- Daily AI Hotspot · 2026-08-26: Adaptive DSpark Delivers +33.6% Throughput at c=256, While Step's MTP Acceptance Rate Silently Collapses from 97% to 4% 2026-08-26T00:00:00.000Z
- Daily AI Hotspot · 2026-08-25: vLLM Ships a ~3× Decode Kernel for 4090D/H20, SGLang Spreads DSpark Verification Across Four Quantization Tiers 2026-08-25T00:00:00.000Z
- Daily AI Hotspot · 2026-08-24: MTP Speculative Decoding Crosses Into Multimodal for the First Time (Nemotron VL); SGLang Quiet 2026-08-24T00:00:00.000Z
- Daily AI Hotspot · 2026-08-23: SGLang v0.5.18 Cuts Startup to 35.6 Seconds, While vLLM v0.28.0 Enters RC with DFlash2 2026-08-23T00:00:00.000Z
- Daily AI Hotspot · 2026-08-21: Shared-Expert Fusion Lands on NVIDIA and AMD the Same Day, +14.98% Throughput at BS=1 2026-08-21T00:00:00.000Z
- Daily AI Hotspot · 2026-08-19: DFlash2 Gives the Draft Head a Local Conv and a Candidate Selector — Speculative Decoding Moves from Acceptance Rate to Draft Quality 2026-08-19T00:00:00.000Z
- Daily AI Hotspot · 2026-08-18: TensorCast Turns KV Migration Into a Shared Layer, Cutting Agent Median TTFT by Up to 93.2% 2026-08-18T00:00:00.000Z
- Daily AI Hotspot · 2026-08-13: Speculative Decoding Starts Doing the Math — Fixed Draft Count Loses 33% at High Concurrency 2026-08-13T00:00:00.000Z