系列:Paper Primer Notes

Transformer Deep Read — Replacing Recurrence with Attention, Making Sequence Modeling Parallel

1. The One-Line Takeaway

The core claim of the Transformer is this: sequence modeling does not need to read “one token at a time” (RNN), nor “through a sliding window” (CNN). Instead, every position directly “looks up” at every other position via a mechanism called self-attention, computing who relates to whom in a single step. This makes it natively parallel, stackable, and scalable to very long contexts — and every LLM you use today is, at its core, a Transformer.

2. Why Replace the RNN?

Before 2017, the workhorses for language, speech, and time series were RNN / LSTM:

CNNs can parallelize, but their receptive field is bounded by kernel size; capturing distant dependencies needs many stacked layers.

The Transformer’s bet: model the dependency between any two positions directly, where distance costs only one extra compute step.

3. Self-Attention: Every Word “Looks at the Whole Field”

Self-attention takes a set of vectors X = [x₁, x₂, …, xₙ] (each token’s embedding). For every position it computes three vectors:

Obtained by linear projection: Q = XW_Q, K = XW_K, V = XW_V. Then for each position, compute relevance to all others:

score(i, j) = Q_i · K_j / √d_k
attention(i) = softmax_j( score(i, j) ) · V_j

Intuition: the more Q_i resembles K_j (larger dot product), the more word i “attends” to word j; the output is the weighted sum of all positions’ V by attention weights.

Role of √d_k: with large d_k, dot products explode. Dividing by its square root stabilizes the softmax gradient — a small but critical trick in the paper.

Key point: any word looking at another is just one matrix multiplication away, regardless of how far apart they are. That is “global dependency in one step”.

4. Multi-Head Attention: Not Just One Angle

A single attention learns only one “relationship view”. The Transformer uses Multi-Head Attention: split Q/K/V into h heads, each computing attention in its own subspace, then concatenate and project:

head_h = Attention(XW_Q^h, XW_K^h, XW_V^h)
MultiHead = Concat(head₁, …, head_h) · W_O

Different heads specialize: some watch syntax, some coreference, some semantics. The paper uses h = 8 heads, each of dimension 64 (total 512).

5. Positional Encoding: Attention Itself Is Order-Agnostic

Self-attention is permutation-invariant — shuffle the sentence and each word’s weighted sum depends only on which words appear, not who comes first. But language has order.

Fix: add a positional encoding PE(pos) to each position:

Later models (e.g. GPT) mostly use learnable positional embeddings — same idea: order information must be fed in.

6. The Encoder–Decoder Skeleton

The original Transformer was built for translation, in two halves:

7. Why It Became the LLM Foundation?

PropertyRNN/LSTMTransformer
ParallelismSerial, near-zeroWhole-sequence matmul, GPU-friendly
Long-rangeLong path, forgetsAny two points, one step
ScalabilityHard to deepenResidual+LayerNorm, 100+ layers
Hardware fitIrregular loopsPure MatMul

Because of these, GPT (decoder-only), BERT (encoder-only), ViT (images-as-sequences), whisper, Sora, and even the action head of Diffusion Policy all sit on this attention mechanism. One 2017 paper defined the “universal compute primitive” of deep learning for the next eight years.

8. Summary & Next Steps

Next (pp2): Diffusion Policy — it ports the other lineage (diffusion denoising, alongside the attention paradigm) into robot action generation, linking neatly with this blog’s “VLA Notes”.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。