系列:Frontier Architecture Decoding Notes

Frontier Architecture Decoding Notes (6): Gated DeltaNet — Conv1D Captures Locality, the Gated Delta Rule Compresses Globality

0. Why a Whole Post for This Layer

Posts fa1 to fa4 covered how to make attention lighter (sparsity, compression, hybrid SSM). Gated DeltaNet (GDN), the subject here, goes further — it is not attention but a linear attention / recurrent state-space layer whose core is a fixed-size “state matrix” updated incrementally at every token, compressing global history into an O(d²) fixed-length tensor.

Its layer structure is unusually clean, with only two operators:

Together they form a “local convolution first, global linear attention second” hierarchy: Conv1D handles short-range dependency and position, while the recurrent state compresses long history into a fixed-length state at O(n·d²). The diagram below is the complete path.

x_t (current token) Q / K / V projection (Linear), no RoPE Position is left to Conv1D; RoPE is not introduced Operator 1 - Causal Conv1D (depthwise separable causal conv + SiLU) Maintains a conv_state sliding window, giving local context plus position awareness Q1 = Conv1D(Q) + SiLU K1 = Conv1D(K) + SiLU V1 = Conv1D(V) + SiLU Operator 2 - Recurrent State Update (Gated Delta Rule) Gates alpha_t, beta_t (data dependent) alpha forgets, beta writes S_t: one per head d_v x d_k matrix (fixed size)

S_t = S_{t-1} * alpha_t ( I - beta_t k_t k_t^T ) + beta_t v_t k_t^T delta-rule write, minus decay gate alpha_t forgetting history; complexity O(n*d^2)

Recurrence: S_t becomes S_{t-1} at the next step, so the fixed state rolls across tokens o_t (layer output)
Figure: the two operators of a Gated DeltaNet layer. Operator 1 applies a depthwise separable causal convolution plus SiLU to each of Q, K and V after projection (maintaining a conv_state sliding window); operator 2 takes the convolved Q1/K1/V1 with data-dependent gates alpha_t and beta_t, applies the Gated Delta Rule to update the fixed-size state matrix S_t, and emits the output. Conv1D handles locality and position; the recurrent state handles global linear attention.

1. Operator 1: Causal Conv1D (Locality Plus Position)

It sits after the Q/K/V projection and does the following separately to Q, K and V:

It serves two purposes:

  1. Short-range dependency modeling. Convolution naturally aggregates neighboring tokens, covering the weakness of linear attention, which “only looks at the global state and is insensitive to locality.”
  2. Position awareness (replacing RoPE). The GDN layer uses no RoPE at all; position information comes from the local receptive field of this causal convolution — so the judgment that “Causal Conv1D acts like a positional encoding or short-range dependency model” is accurate.

2. Operator 2: Recurrent State Update (Gated Delta Rule)

This is the core of GDN. It takes the convolved Q’/K’/V’ along with two data-dependent gate signals:

Complete formula (including the decay gate alpha_t):

S_t = S_{t-1} * alpha_t ( I - beta_t k_t k_t^T )  +  beta_t v_t k_t^T

Term by term:

Note: the form S_t = S_{t-1} - beta_t (S_{t-1} phi(k_t)) phi(k_t)^T + beta_t v_t phi(k_t)^T that appears in some documentation is the naive delta rule without the alpha_t decay gate, matching the original paper; but the complete GDN formula must include the alpha_t decay term, otherwise it degenerates into linear attention with no forgetting mechanism.

3. Positional Encoding: No RoPE

The GDN layer does not use rotary position embedding (RoPE). All absolute and relative position cues in the layer come from:

That is precisely why GDN can keep linear complexity without leaning heavily on injected positional encodings.

4. State Shape: A Matrix, Not a Vector

ssm_state is shaped as one d_v x d_k matrix per attention head (not a single vector):

The “state matrix (or vector)” wording in some documentation is vague; it should be stated plainly as a matrix: each head holds a fixed-size d_v x d_k table on which the recurrence performs “forget, correct, write.”

5. Complexity: Why It Is Linear

When the sequence is long (n ≫ d), O(n·d²) is dramatically lower than O(n²·d). That is the fundamental reason GDN sustains long context without a KV cache that explodes with sequence length — its “memory” is a fixed-size matrix, not a per-token KV list.

6. Relationship to Qwen3.5 / 3.6 / 3.8 (Model Name Correction)

The exact model name qwen3.8-27B appearing in the documentation is now confirmed to exist. An earlier judgment in this post that it “does not exist” was wrong and has been corrected against a 2026-08-28 repository snapshot scan (see post 7 for cross-validation of the Qwen3.8 dual checkpoints). The main Qwen3.8 tier is a MoE model such as 2.4T-A95B, and a 27B dense checkpoint also exists.

One thing, however, is certain: the Gated DeltaNet architecture (the GDN layer) really is a core component of the Qwen3.5 / 3.6 / 3.8 family, and a key module behind that family’s linear inference cost at long context. Therefore:

7. Four Corrections After Cross-Checking the Documentation

Cross-checking the original paper (Yang et al., Gated DeltaNet, ICLR 2025) against the Qwen3.5/3.6/3.8 implementation code, the overall description is essentially correct, with four points to fix:

#PointDocumentation saidCorrect statement
1Model nameqwen3.8-27BThe 27B dense checkpoint does exist (confirmed by a 2026-08-28 repository scan; the earlier “does not exist” judgment is withdrawn); GDN is a core component of Qwen3.5/3.6/3.8 ✅
2Complete formulaDelta rule with beta onlyMust add the decay gate alpha: S_t = S_{t-1} * alpha_t (I - beta_t k_t k_t^T) + beta_t v_t k_t^T
3Positional encodingA passing mention that Conv1D is “like a positional encoding”Accurate — GDN has no RoPE at all; position comes from the causal convolution plus state rolling
4State shape“State matrix (or vector)”, slightly vaguePrecisely one d_v x d_k matrix per head (fixed size, independent of sequence length)
One line to remember: Gated DeltaNet = Conv1D for locality and position plus the Gated Delta Rule for global compression; the state is one fixed-size d_v x d_k matrix per head, forgetting with alpha and writing with beta, with no RoPE anywhere and O(n*d^2) complexity. It is one of the pillars of linear inference cost in the Qwen3.5/3.6/3.8 family.

Next up: a horizontal comparison of GDN with Mamba and the linear attention family (GLA, RWKV, RetNet) — all of them “replace the KV list with a fixed-size state,” but their state update rules are completely different.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。