大模型技术notes

Today · Autumn Equinox · Sunday, September 27, 2026

Notes on LLM tech and multimodality. Chinese-first; twelve key series (Speculative Decoding Notes, Multimodal Decoding Notes, VLA Decoding Notes, vLLM & SGLang Serving Notes, Frontier Architecture Decoding Notes, Inference Systems Infrastructure Notes, RL Training Notes, Agent Notes, Operator Notes, Paper Primer Notes, Community Tracker Notes, Daily AI Hotspot Notes) keep growing, with most posts bilingual.

Speculative Decoding Notes

Series · 8 episodes, in order: fundamentals → the autoregressive drafter's ceiling → DFlash / DSpark deep dives → a head-to-head → EAGLE-3 (why the feature constraint is a scaling lock) → MTP & confidence heads (three self-draft routes chained) → DFlash 2 upgrade and the meta-question of why frameworks don't just combine them. Each post has a Chinese counterpart.

Multimodal Decoding Notes

Series · 5 episodes, in order: foundations & alignment → VLM evolution → generative models (diffusion / GAN) → VLA & world models → efficient systems & frontier inference. A progressive arc through the common multimodal models, structures, principles, and optimization techniques, each with a Chinese counterpart.

vLLM & SGLang Serving Notes

Series · 7 episodes, in order: why an inference engine → vLLM internals → SGLang internals → the shared frontier (sync stalls / spec decode / PD disagg / low-bit quant) → the July 2026 release waves → supported models & selection guide → graph mode (CUDA Graphs). The daily vLLM/SGLang tracking distilled into one arc, with hand-drawn SVG architecture diagrams, each with a Chinese counterpart.

Inference Systems Infrastructure Notes

Series · 11 episodes, in order: the NUMA x PCIe x NIC data path under P900 + HiCache → Sequence Parallelism + Ring Attention for 1M+ token training → MSA / CSA / HCA as three attention-redesign routes → knowledge distillation and dark knowledge → the KV cache compression landscape from MHA to HiCache → HiSparse hierarchical KV cache (treating HBM as a cache to break the capacity wall) → DCP decode context parallelism (sharding KV along the sequence dimension, reusing a Scratch Buffer for comms) → parallelism strategies and PD disaggregation (how to slice and combine TP/PP/EP/SP/CP, and how to compute the prefill:decode instance ratio) → the KV Cache landscape (why it can be cached, how big it gets, the three-layer taxonomy and five quantifiable knobs) → KV Cache in production (SGLang two-level pools and HiCache, vLLM paged block management and the cross-layer unified layout, and AttentionStore measured on Kunlunxin P800) → tuning SGLang HiCache L2 on Kunlunxin P800/P900 (tiered architecture and net-benefit formula → layout choice page_first_direct to layer_first and DMA segment count → async transfer, NUMA affinity, transparent huge pages, PCIe IDO ordering → write-policy semantic fix and Ghost List, TP group shared Host KV → end-to-end TTFT from 28.83s to 18.45s). It answers "once the model has computed, what does the systems layer still owe you," and pairs with the model-side redesign in Frontier Architecture Decoding Notes.

Agent Notes

Series · 1 episode (opening, ongoing). Threading the scattered Agent notes into one system: capability stack (model / planning / tools / memory) → tool calling and MCP → context engineering (1M-token directory attention) → execution paradigms and orchestration (Orchestrator + executors) → how Agentic workloads turn KV Cache from a memory problem into a scheduling problem. The core claim: Agents define a variable-length, branching, multi-turn workload shape that breaks the engine static scheduling assumptions, closing the loop with HiCache / HiSparse / DCP / KV quantization in Inference Systems Infrastructure Notes. Each post has a Chinese counterpart.

VLA Decoding Notes

Series · 9 episodes, in order: from "see" to "act" (what VLA is) → the action engine (Diffusion Policy / Flow Matching / Action Chunking / low latency) → the π family (π0→π0.7) → domestic players (Ant Lingbo / Xiaomi / Tencent) → world models & deployment (data flywheel / scaling bottleneck / investment view) → RT-1's discrete-action route (TokenLearner / causal mask / cross-entropy) → V-JEPA 2 self-supervised video world models → Dreamer V3 world-model RL → Octo (a 27M generalist robot policy that turns fine-tuning to new embodiments into an engineering problem). A natural extension of the LLaVA piece in Multimodal Decoding Notes (2) into the robot's "act" capability, each with a Chinese counterpart.

Frontier Architecture Decoding Notes

Series · 8 episodes, in order: the 2026 frontier overview (the impossible triangle of long-context × efficiency × scale) → Kimi K3 (redesign attention: KDA + Gated MLA 3:1 + Stable LatentMoE + native vision) → MiniMax M2.7→M3 (sparse attention MSA cuts 1M-context compute to 1/20) → DeepSeek V4 (compress attention CSA+HCA + Muon optimizer + OPD distillation, KV down to ~10%) → Mamba hybrid state + PD disaggregation (how the KV plus SSM recurrent state of Jamba-style hybrids crosses nodes, and why the Mamba state has no token-id key and silently mis-hits) → Gated DeltaNet (Conv1D captures locality, the gated delta rule compresses globality) → Qwen3.8 dual checkpoints (dense 27B vs sparse Flash-Next, 360GB of weights for 6B activations) → the Qwen3.8 family in one frame (2.4T-A95B flagship / Flash serving build / Flash-Next architecture preview, 125B MoE + 51B N-gram at 6B active per token). The attention-reshape controlled experiments strung into one arc: dense → sparse → compressed attention → linear attention and hybrid state → sparse MoE at scale, each with a Chinese counterpart.

Operator Notes

Series · 2 episodes (ongoing). Turn the operators frontier architectures gloss over into line-by-line pseudocode you can implement: episode one is MLA's c_kv / W_DKV / W_UK / W_UV operator and its cache-compression magnitude; episode two is model quantization for deployment (FP8/INT4/AWQ·GPTQ/FP4 QAT/KV Cache quantization); later posts cover GQA / MQA, normalization, routing and more. Each post has a Chinese counterpart.

Paper Primer Notes

Series · 2 episodes (opening). Readable notes on the "classic papers" no one in the LLM era can avoid: from Transformer (2017), which defined attention, to Diffusion Policy (2023), which ports the diffusion-denoising paradigm into robot action generation. Complements the "action engine" in VLA Notes and the engineering series; each post has a Chinese counterpart.

Community Tracker Notes

Series · 3 episodes. Daily vLLM / SGLang upstream commits, model cookbooks, and domestic-model enablement progress, framed as "same-day same-frame / who is falling behind" — filling the timeliness gap left by the principle-heavy framework notes; each post has a Chinese counterpart.

RL Training Notes

Series · 1 episode (opening, ongoing). Filling a gap the site previously had: the training-side infrastructure shared by LLMs and robots. The opening post covers distributed RL training — why RLHF cannot run as a single-process pot of everything (rollout eats roughly 80% of a PPO step), how the Actor / Critic / Reference / Reward roles decouple onto separate GPU pools, how vLLM or SGLang serves as the rollout engine, the PPO and GRPO objectives and the price of deleting the Critic, the two weight-sync paths (Ray Object Store vs NCCL broadcast), and how the same stack serves both text RLHF and robot VLA RLAIF. Complements the inference-side optimization in Inference Systems Infrastructure Notes, each post with a Chinese counterpart.

Daily AI Hotspot Notes

Series · 27 episodes (ongoing). Front-line daily progress on vLLM / SGLang upstream, domestic-model enablement, speculative decoding and inference optimization, framed as "same-day same-frame / who is falling behind" — filling the timeliness gap left by the principle-heavy framework notes; each post has a Chinese counterpart.

View all 27 episodes →

Found these notes useful? Buy the author a coffee ☕️

Alipay

Alipay

WeChat Pay

WeChat