大模型技术notes
Notes on LLM tech and multimodality. Chinese-first; twelve key series (Speculative Decoding Notes, Multimodal Decoding Notes, VLA Decoding Notes, vLLM & SGLang Serving Notes, Frontier Architecture Decoding Notes, Inference Systems Infrastructure Notes, RL Training Notes, Agent Notes, Operator Notes, Paper Primer Notes, Community Tracker Notes, Daily AI Hotspot Notes) keep growing, with most posts bilingual.
Speculative Decoding Notes
Series · 8 episodes, in order: fundamentals → the autoregressive drafter's ceiling → DFlash / DSpark deep dives → a head-to-head → EAGLE-3 (why the feature constraint is a scaling lock) → MTP & confidence heads (three self-draft routes chained) → DFlash 2 upgrade and the meta-question of why frameworks don't just combine them. Each post has a Chinese counterpart.
- Speculative Decoding: Lossless Multi-Token Generation 2026-07-31
- The Autoregressive Drafter's Ceiling: Why It Stalls at 2–3× 2026-07-31
- DFlash Deep Dive: Block-Diffusion Drafting, KV Injection, and a Bubble-Free Pipeline 2026-07-31
- DSpark Deep Dive: Semi-Autoregressive Drafting with a Confidence Scheduler 2026-07-31
- Head-to-Head: Where DFlash and DSpark Actually Differ 2026-07-31
- EAGLE-3 Deep Dive: Drop the Feature Constraint, Let the Drafter Scale 2026-08-04
- MTP Heads and Confidence Heads: Linking Three Self-Drafting Routes into One Chain 2026-08-19
- DFlash 2 and the Combine Question: Block-Diffusion Upgrade, but Why Not Stack DSpark Scheduling? 2026-09-09
Multimodal Decoding Notes
Series · 5 episodes, in order: foundations & alignment → VLM evolution → generative models (diffusion / GAN) → VLA & world models → efficient systems & frontier inference. A progressive arc through the common multimodal models, structures, principles, and optimization techniques, each with a Chinese counterpart.
- (1) Foundations & Alignment — from "Writing" to "Seeing & Acting" 2026-07-31
- (2) VLM Evolution — ViT → CLIP → LLaVA 2026-07-31
- (3) Generative Models — Diffusion vs GAN (Stable Diffusion & GAN) 2026-07-31
- (4) From VLA to World Models — RT-2 → π0.7 → HiF-VLA 2026-07-31
- (5) Efficient Systems & Frontier Inference — Attention / Inference Opt / On-Device / RL 2026-07-31
vLLM & SGLang Serving Notes
Series · 7 episodes, in order: why an inference engine → vLLM internals → SGLang internals → the shared frontier (sync stalls / spec decode / PD disagg / low-bit quant) → the July 2026 release waves → supported models & selection guide → graph mode (CUDA Graphs). The daily vLLM/SGLang tracking distilled into one arc, with hand-drawn SVG architecture diagrams, each with a Chinese counterpart.
- (1) Why an Inference Engine — Three Walls of Naive Serving 2026-08-01
- (2) vLLM Internals — From Paged KV to Model Runner V2 2026-08-01
- (3) SGLang Internals — Prefix Reuse & Frontier Throughput 2026-08-01
- (4) The Shared Frontier — Sync Stalls, Spec Decode, PD, Low-Bit Quant 2026-08-01
- (5) The July 2026 Release Waves — vLLM 0.25→0.26 & SGLang 0.5.15→0.5.16 2026-08-01
- (6) Supported Models & Selection Guide — Who to Pick 2026-08-01
- (7) Graph Mode — CUDA Graphs in vLLM & SGLang 2026-08-04
- (8) Two Layers of V1/V2 in vLLM 0.25.1 & Where vLLM-Kunlun Lands 2026-08-05
- (9) Wan2.2-I2V/T2V-A14B: Dual-Expert MoE Video Model and vllm-Omni Engine Dissected 2026-08-11
Inference Systems Infrastructure Notes
Series · 11 episodes, in order: the NUMA x PCIe x NIC data path under P900 + HiCache → Sequence Parallelism + Ring Attention for 1M+ token training → MSA / CSA / HCA as three attention-redesign routes → knowledge distillation and dark knowledge → the KV cache compression landscape from MHA to HiCache → HiSparse hierarchical KV cache (treating HBM as a cache to break the capacity wall) → DCP decode context parallelism (sharding KV along the sequence dimension, reusing a Scratch Buffer for comms) → parallelism strategies and PD disaggregation (how to slice and combine TP/PP/EP/SP/CP, and how to compute the prefill:decode instance ratio) → the KV Cache landscape (why it can be cached, how big it gets, the three-layer taxonomy and five quantifiable knobs) → KV Cache in production (SGLang two-level pools and HiCache, vLLM paged block management and the cross-layer unified layout, and AttentionStore measured on Kunlunxin P800) → tuning SGLang HiCache L2 on Kunlunxin P800/P900 (tiered architecture and net-benefit formula → layout choice page_first_direct to layer_first and DMA segment count → async transfer, NUMA affinity, transparent huge pages, PCIe IDO ordering → write-policy semantic fix and Ghost List, TP group shared Host KV → end-to-end TTFT from 28.83s to 18.45s). It answers "once the model has computed, what does the systems layer still owe you," and pairs with the model-side redesign in Frontier Architecture Decoding Notes.
- (1) The NUMA x PCIe x NIC Data Path Under P900 + HiCache2026-08-26
- (2) Sequence Parallelism + Ring Attention: Making 1M+ Token Training Routine2026-08-26
- (3) MSA / CSA / HCA: Three Attention Redesign Routes in One Picture2026-08-26
- (4) Knowledge Distillation: Passing the Dark Knowledge of a Large Model to a Small One2026-08-26
- (5) From MHA to HiCache: The Full KV Cache Compression Landscape2026-08-26
- (6) HiSparse: Treating HBM as a Cache to Break the Long-Context Capacity Wall2026-09-07
- (7) DCP: Decode Context Parallelism — Sharding KV Cache Along the Sequence Dimension2026-09-08
- (8) Parallelism Strategies and PD Disaggregation — slicing TP/PP/EP/SP/CP and sizing the prefill:decode ratio2026-09-08
- (9) The KV Cache Landscape — why it decides your throughput, latency and concurrency2026-09-09
- (10) KV Cache in Production — SGLang two-level pools, vLLM block management, and AttentionStore tiered caching2026-09-09
- (11) Tuning SGLang HiCache L2 on Kunlunxin P800/P900: Layout, DMA, NUMA and TP-Shared KV2026-09-15
Agent Notes
Series · 1 episode (opening, ongoing). Threading the scattered Agent notes into one system: capability stack (model / planning / tools / memory) → tool calling and MCP → context engineering (1M-token directory attention) → execution paradigms and orchestration (Orchestrator + executors) → how Agentic workloads turn KV Cache from a memory problem into a scheduling problem. The core claim: Agents define a variable-length, branching, multi-turn workload shape that breaks the engine static scheduling assumptions, closing the loop with HiCache / HiSparse / DCP / KV quantization in Inference Systems Infrastructure Notes. Each post has a Chinese counterpart.
VLA Decoding Notes
Series · 9 episodes, in order: from "see" to "act" (what VLA is) → the action engine (Diffusion Policy / Flow Matching / Action Chunking / low latency) → the π family (π0→π0.7) → domestic players (Ant Lingbo / Xiaomi / Tencent) → world models & deployment (data flywheel / scaling bottleneck / investment view) → RT-1's discrete-action route (TokenLearner / causal mask / cross-entropy) → V-JEPA 2 self-supervised video world models → Dreamer V3 world-model RL → Octo (a 27M generalist robot policy that turns fine-tuning to new embodiments into an engineering problem). A natural extension of the LLaVA piece in Multimodal Decoding Notes (2) into the robot's "act" capability, each with a Chinese counterpart.
- (1) From "See" to "Act" — What Is a Vision-Language-Action Model 2026-08-04
- (2) The Action Engine — Diffusion Policy, Flow Matching, Low Latency 2026-08-04
- (3) The π Family — Physical Intelligence’s VLA Lineage 2026-08-04
- (4) Domestic Players — Ant Lingbo, Xiaomi, Tencent 2026-08-04
- (5) World Models and Deployment — VLA’s Next Stop 2026-08-04
- (6) RT-1's Discrete-Action Route — TokenLearner, Causal Mask, and Cross-Entropy 2026-08-07
- (7) V-JEPA 2: Self-Supervised Video World Models and Zero-Shot Robot Planning 2026-08-26
- (8) Dreamer V3: One Set of Hyperparameters Across 50+ Domains 2026-08-26
- (9) Octo: A 27M Generalist Robot Policy and What Fine-Tunable Really Means 2026-09-11
Frontier Architecture Decoding Notes
Series · 8 episodes, in order: the 2026 frontier overview (the impossible triangle of long-context × efficiency × scale) → Kimi K3 (redesign attention: KDA + Gated MLA 3:1 + Stable LatentMoE + native vision) → MiniMax M2.7→M3 (sparse attention MSA cuts 1M-context compute to 1/20) → DeepSeek V4 (compress attention CSA+HCA + Muon optimizer + OPD distillation, KV down to ~10%) → Mamba hybrid state + PD disaggregation (how the KV plus SSM recurrent state of Jamba-style hybrids crosses nodes, and why the Mamba state has no token-id key and silently mis-hits) → Gated DeltaNet (Conv1D captures locality, the gated delta rule compresses globality) → Qwen3.8 dual checkpoints (dense 27B vs sparse Flash-Next, 360GB of weights for 6B activations) → the Qwen3.8 family in one frame (2.4T-A95B flagship / Flash serving build / Flash-Next architecture preview, 125B MoE + 51B N-gram at 6B active per token). The attention-reshape controlled experiments strung into one arc: dense → sparse → compressed attention → linear attention and hybrid state → sparse MoE at scale, each with a Chinese counterpart.
- (1) The 2026 Frontier Overview — the Impossible Triangle of Long-Context × Efficiency × Scale 2026-08-06
- (2) Kimi K3 — Dual Evolution of Attention and MoE 2026-08-06
- (3) MiniMax M2.7 → M3 — The Leap of Sparse Attention 2026-08-06
- (4) DeepSeek V4 — Extreme Compression and Efficient Training 2026-08-06
- (5) Mamba Hybrid State + PD Disaggregation — Why Cache Correctness Must Avoid Silent Hit Misplacement 2026-08-10
- (6) Gated DeltaNet — Conv1D Captures Locality, the Gated Delta Rule Compresses Globality 2026-08-26
- (7) Qwen3.8 Dual Checkpoints — Dense 27B vs Sparse Flash-Next (360GB of Weights for 6B Activations) 2026-08-28
- (8) The Qwen3.8 Family in One Frame — the 2.4T Flagship, Flash the Serving Build, and Flash-Next the Architecture Preview 2026-09-09
Operator Notes
Series · 2 episodes (ongoing). Turn the operators frontier architectures gloss over into line-by-line pseudocode you can implement: episode one is MLA's c_kv / W_DKV / W_UK / W_UV operator and its cache-compression magnitude; episode two is model quantization for deployment (FP8/INT4/AWQ·GPTQ/FP4 QAT/KV Cache quantization); later posts cover GQA / MQA, normalization, routing and more. Each post has a Chinese counterpart.
- (1) MLA Operator-Level — c_kv / W_DKV / W_UK / W_UV Line-by-Line Pseudocode 2026-08-07
- (2) Model Quantization for Deployment — the precision-compression spectrum from FP16 to FP4 2026-09-08
Paper Primer Notes
Series · 2 episodes (opening). Readable notes on the "classic papers" no one in the LLM era can avoid: from Transformer (2017), which defined attention, to Diffusion Policy (2023), which ports the diffusion-denoising paradigm into robot action generation. Complements the "action engine" in VLA Notes and the engineering series; each post has a Chinese counterpart.
- (1) Transformer Deep Read — Replacing Recurrence with Attention, Making Sequence Modeling Parallel 2026-08-12
- (2) Diffusion Policy Deep Read — Porting Stable Diffusion’s Denoising to Robot Action Generation 2026-08-12
Community Tracker Notes
Series · 3 episodes. Daily vLLM / SGLang upstream commits, model cookbooks, and domestic-model enablement progress, framed as "same-day same-frame / who is falling behind" — filling the timeliness gap left by the principle-heavy framework notes; each post has a Chinese counterpart.
- vLLM & SGLang Community Tracker · 2026-08-06 — Both Frameworks Land Ling-3.0-flash Same Day; Step Stalls 2026-08-12
- vLLM & SGLang Community Tracker · 2026-08-07 — KV Reuse Enters "Combinatorial Correctness" (HiCache / Mooncake) 2026-08-12
- vLLM & SGLang Community Tracker · 2026-08-11 — vLLM v0.27.0 Ships in One Drop; SGLang Triple-Fixes Unified Pool & Wires PD 2026-08-12
RL Training Notes
Series · 1 episode (opening, ongoing). Filling a gap the site previously had: the training-side infrastructure shared by LLMs and robots. The opening post covers distributed RL training — why RLHF cannot run as a single-process pot of everything (rollout eats roughly 80% of a PPO step), how the Actor / Critic / Reference / Reward roles decouple onto separate GPU pools, how vLLM or SGLang serves as the rollout engine, the PPO and GRPO objectives and the price of deleting the Critic, the two weight-sync paths (Ray Object Store vs NCCL broadcast), and how the same stack serves both text RLHF and robot VLA RLAIF. Complements the inference-side optimization in Inference Systems Infrastructure Notes, each post with a Chinese counterpart.
Daily AI Hotspot Notes
Series · 27 episodes (ongoing). Front-line daily progress on vLLM / SGLang upstream, domestic-model enablement, speculative decoding and inference optimization, framed as "same-day same-frame / who is falling behind" — filling the timeliness gap left by the principle-heavy framework notes; each post has a Chinese counterpart.
- Daily AI Hotspot · 2026-09-27: No New Release on Sunday — SGLang main Closes Cache/Routing Base, kv-hints Envelope Enters Transport; DriveVLM Puts VLM Slow-Thinking Into the Autonomous-Driving Loop 2026-09-27T00:00:00.000Z
Found these notes useful? Buy the author a coffee ☕️
Alipay