0. Why a Whole Post for This Layer
Posts fa1 to fa4 covered how to make attention lighter (sparsity, compression, hybrid SSM). Gated DeltaNet (GDN), the subject here, goes further — it is not attention but a linear attention / recurrent state-space layer whose core is a fixed-size “state matrix” updated incrementally at every token, compressing global history into an O(d²) fixed-length tensor.
Its layer structure is unusually clean, with only two operators:
- Operator 1, Causal Conv1D: after the Q/K/V projections, each of the three branches gets a depthwise separable causal convolution plus SiLU, providing local context and position awareness.
- Operator 2, Recurrent State Update: takes the convolved Q’/K’/V’ plus data-dependent decay and write gates alpha_t and beta_t, applies the Gated Delta Rule, and updates the fixed-size
ssm_statematrix to produce the output.
Together they form a “local convolution first, global linear attention second” hierarchy: Conv1D handles short-range dependency and position, while the recurrent state compresses long history into a fixed-length state at O(n·d²). The diagram below is the complete path.
1. Operator 1: Causal Conv1D (Locality Plus Position)
It sits after the Q/K/V projection and does the following separately to Q, K and V:
- Depthwise separable causal Conv1D: a one-dimensional causal convolution applied independently per channel; the kernel slides only over the current and past positions, preserving causality (no peeking at the future).
- SiLU activation after the convolution.
- conv_state sliding window: a fixed-length sliding-window state is maintained during training and inference, so recurrent stepping only needs to cache the most recent
kernel_sizepositions — it does not grow with the sequence.
It serves two purposes:
- Short-range dependency modeling. Convolution naturally aggregates neighboring tokens, covering the weakness of linear attention, which “only looks at the global state and is insensitive to locality.”
- Position awareness (replacing RoPE). The GDN layer uses no RoPE at all; position information comes from the local receptive field of this causal convolution — so the judgment that “Causal Conv1D acts like a positional encoding or short-range dependency model” is accurate.
2. Operator 2: Recurrent State Update (Gated Delta Rule)
This is the core of GDN. It takes the convolved Q’/K’/V’ along with two data-dependent gate signals:
- alpha_t (decay gate, in (0,1)): controls how strongly the historical state is exponentially forgotten.
- beta_t (write gate): controls how strongly the current token is written into the state.
Complete formula (including the decay gate alpha_t):
S_t = S_{t-1} * alpha_t ( I - beta_t k_t k_t^T ) + beta_t v_t k_t^T
Term by term:
S_{t-1} * alpha_t: the previous state is first scaled by the decay gate as a whole, so history is forgotten in proportion to alpha_t (this is exactly the term the naive delta rule lacks).alpha_t ( I - beta_t k_t k_t^T ): on top of the decayed state, a “delta correction along the current key direction” is applied —beta_t k_t k_t^Tsubtracts from the old state the projection of “the old prediction associated with the current key,” freeing up capacity.+ beta_t v_t k_t^T: the current value is written into the state scaled by the write gate beta_t, as the outer product of “key to value.”
Note: the form
S_t = S_{t-1} - beta_t (S_{t-1} phi(k_t)) phi(k_t)^T + beta_t v_t phi(k_t)^Tthat appears in some documentation is the naive delta rule without the alpha_t decay gate, matching the original paper; but the complete GDN formula must include the alpha_t decay term, otherwise it degenerates into linear attention with no forgetting mechanism.
3. Positional Encoding: No RoPE
The GDN layer does not use rotary position embedding (RoPE). All absolute and relative position cues in the layer come from:
- the causal convolution receptive field of operator 1 (local ordering);
- the recurrent state
S_titself rolling forward in time order (implicit temporal structure).
That is precisely why GDN can keep linear complexity without leaning heavily on injected positional encodings.
4. State Shape: A Matrix, Not a Vector
ssm_state is shaped as one d_v x d_k matrix per attention head (not a single vector):
d_kis the key dimension (one side of the outer productk_t k_t^T);d_vis the value dimension (the other side of the outer productv_t k_t^T);- the matrix size is independent of sequence length, so the state stays constant in size — this is the foundation of O(n·d²) linear complexity.
The “state matrix (or vector)” wording in some documentation is vague; it should be stated plainly as a matrix: each head holds a fixed-size d_v x d_k table on which the recurrence performs “forget, correct, write.”
5. Complexity: Why It Is Linear
- Naive attention: every token takes a dot product against the entire history, complexity O(n²·d).
- GDN recurrent state: every token only does “read state, update a fixed-size matrix, write state,” with state size fixed at
d_v·d_k, complexity O(n·d²).
When the sequence is long (n ≫ d), O(n·d²) is dramatically lower than O(n²·d). That is the fundamental reason GDN sustains long context without a KV cache that explodes with sequence length — its “memory” is a fixed-size matrix, not a per-token KV list.
6. Relationship to Qwen3.5 / 3.6 / 3.8 (Model Name Correction)
The exact model name qwen3.8-27B appearing in the documentation is now confirmed to exist. An earlier judgment in this post that it “does not exist” was wrong and has been corrected against a 2026-08-28 repository snapshot scan (see post 7 for cross-validation of the Qwen3.8 dual checkpoints). The main Qwen3.8 tier is a MoE model such as 2.4T-A95B, and a 27B dense checkpoint also exists.
One thing, however, is certain: the Gated DeltaNet architecture (the GDN layer) really is a core component of the Qwen3.5 / 3.6 / 3.8 family, and a key module behind that family’s linear inference cost at long context. Therefore:
- Correction: “Qwen3.8-27B does exist” (a 27B dense checkpoint, confirmed by a 2026-08-28 repository scan; the earlier “does not exist” judgment is withdrawn).
- Correct: “GDN is a core architecture component of the Qwen3.5/3.6/3.8 family (including the 2.4T-A95B MoE and the 27B dense variant).”
7. Four Corrections After Cross-Checking the Documentation
Cross-checking the original paper (Yang et al., Gated DeltaNet, ICLR 2025) against the Qwen3.5/3.6/3.8 implementation code, the overall description is essentially correct, with four points to fix:
| # | Point | Documentation said | Correct statement |
|---|---|---|---|
| 1 | Model name | qwen3.8-27B | The 27B dense checkpoint does exist (confirmed by a 2026-08-28 repository scan; the earlier “does not exist” judgment is withdrawn); GDN is a core component of Qwen3.5/3.6/3.8 ✅ |
| 2 | Complete formula | Delta rule with beta only | Must add the decay gate alpha: S_t = S_{t-1} * alpha_t (I - beta_t k_t k_t^T) + beta_t v_t k_t^T |
| 3 | Positional encoding | A passing mention that Conv1D is “like a positional encoding” | Accurate — GDN has no RoPE at all; position comes from the causal convolution plus state rolling |
| 4 | State shape | “State matrix (or vector)”, slightly vague | Precisely one d_v x d_k matrix per head (fixed size, independent of sequence length) |
Next up: a horizontal comparison of GDN with Mamba and the linear attention family (GLA, RWKV, RetNet) — all of them “replace the KV list with a fixed-size state,” but their state update rules are completely different.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。