K3 Architecture Illustrated
Start with the computation of a single token.
Understand how experts, attention, and depth-wise information flow evolved.
K3 brings several research directions together. KDA and MLA handle sequence history; Stable LatentMoE activates selected feed-forward parameters on demand; AttnRes organizes information flow through a deep network. To understand their roles, first distinguish positions, experts, and layers as three separate concepts.[1] [4]
The big picture: locating Chapters 1–9
From the text backbone to one layer, then to attention, experts, and depth connections.
Follow the full model across the top, then examine the expanded layer below. Chapter 1 provides the layer structure and Chapter 3 its depth connections. Chapters 2 and 4–6 expand the feed-forward path; Chapters 7–9 expand sequence mixing. Click a chapter label in the diagram to jump to it.
Text tokens pass through embeddings, 93 complete layers, final depth aggregation and normalization, and a vocabulary projection to obtain the next-token distribution. Chapter 8 locates KDA and MLA across layers. The lower panel expands a within-block MoE layer: Chapter 1 locates sequence mixing before feed-forward computation, and Chapter 3 supplies each sublayer with its own depth read at the same token position. Within sequence mixing, Chapter 7 covers MLA caches and MHA/GQA/MQA comparisons, while Chapter 9 covers order information and NoPE in later MLA layers. Within feed-forward computation, Chapter 2 locates expert routing and aggregation, Chapter 6 adjusts selection bias, Chapter 5 places routed experts between down- and up-projections, and Chapter 4 locates the gating function inside each expert. The shared branch stays at full width. Detailed normalization and residual-accumulation wiring is omitted.
Inside a Transformer layer
Start with the complete module, then distinguish layers, sublayers, and depth blocks.
Identify the sources and their roles: historical positions, earlier depths, and feed-forward transforms. FIG 01 / Information sources Where does a token get its information? Identify the sources and their roles: historical positions, earlier depths, and feed-forward transforms. States Experts Depth Compute Earlier depths, same token Earlier block b₁ Earlier block b₂ Earlier block b₃ Retrieve earlier states AttnRes Historical token sequence Read historical positions x₁ x₂ x₃ … xₜ₋₁ Current token xₜ Routed MoE experts Selected FFNs transform it E1 E2 E3 E4 E5 E6 E7 E8 Return transformed features Experts perform no cross-token attention Horizontal: sequence position Three paths affect one state Read positions on the left, depths above, and expert transforms on the right. The three paths differ in mechanism; arrows do not imply a simultaneous three-way sum. Source: Kimi K3 report, §2.1–2.3. Expert counts are illustrative; see FIG 02A for placement within a layer.
For a sequence of T tokens, each represented by d numbers, the input forms a T×d table. Each row is a token; each column is a feature coordinate. A layer receives this table, updates the features in each row, and returns another T×d table. The sequence-mixing sublayer moves information across positions; the feed-forward sublayer transforms features independently at each position. Normalization and depth-wise information paths connect these operations.[6] [3]
FIG 02: sequence mixing, feed-forward computation, and depth retrieval within a block. The first depth read combines saved earlier blocks with the current running sum, then normalizes the result for KDA or MLA. A gray bypass retains the prior sum and adds the new sublayer output. The second depth read combines the updated sum with earlier blocks, then normalizes the result for MoE. Its output is again added to the prior sum. Depth retrieval keeps the token position fixed. Other tokens are also processed; their paths are omitted. At a block boundary, the completed block is saved before a new sum begins. The first layer uses a dense FFN.
Read the dataflow together with its coordinates. Moving across the rectangles changes token position. Moving downward through layers changes processing depth. Looking inside one rectangle reveals its feature coordinates.
Across a layer: positions. Down one position: depth. Inside one rectangle: feature values. FIG 02B / Three dimensions Three dimensions: position, depth, and features Across a layer: positions. Down one position: depth. Inside one rectangle: feature values. States Experts Depth Compute Token sequence Position 1 Position 2 Position 3 Position t Earlier layer x₁ x₂ x₃ xₜ Middle layer x₁ x₂ x₃ xₜ Deeper layer x₁ x₂ x₃ xₜ Current layer x₁ x₂ x₃ xₜ Network depth Fix position t Enlarge the current token's vector xₜ Inner cells = feature coordinates All belong to one token. Eight illustrative coordinates, not eight separate tokens. Horizontal is position, vertical is depth, and the vector's interior is feature space. A blue rectangle is a complete token representation; orange marks earlier depths at the same position. Dimension schematic: the number of cells does not represent K3's hidden width.
X_next = X′ + FFN(Norm(X′))
The count of 93 layers refers to complete Transformer layers. Each complete K3 layer contains one sequence-mixing sublayer and one feed-forward sublayer. KDA versus MLA specifies the first; dense FFN versus MoE specifies the second. Both descriptions therefore apply to each layer index.[2] [3]
| Term | Meaning | How it is counted in K3 |
|---|---|---|
| Transformer layer | One sequence-mixing sublayer and one feed-forward sublayer, with their normalization and depth connections | 93 layers in total |
| MoE expert | One set of trainable matrices inside a feed-forward sublayer | Each MoE sublayer has its own expert pool |
| AttnRes depth block | A summary formed by accumulating sublayer outputs across several complete layers | 12 complete layers per block; 9 in the final block |
MoE: a few feed-forward networks per token
Start with a small 2-of-8 example before considering 16 of 896.
A conventional feed-forward network applies the same matrices to every token. A Mixture of Experts (MoE) provides several feed-forward networks plus a router. The router scores the experts from the current token's representation, selects a few to execute, and combines their outputs by weight. This chapter considers sparse MoE with token-choice routing.[14] [15]
Different tokens can call different feed-forward subnetworks: computation is selected by the input. FIG 03 / Conditional computation MoE: more total capacity, less active work per token Different tokens can call different feed-forward subnetworks: computation is selected by the input. States Experts Depth Compute More capacity, less work per token Dense feed-forward network xₜ One FFN for this layer Every token uses the full set of FFN weights MoE / sparse feed-forward experts xₜ Select Expand the pool; use only a few experts per token Total capacity and active parameters can scale separately. Specialization is learned, not assigned by hand. 01 Score xₜ Router 1 2 0.9 3 4 5 6 7 0.6 8 Example scores for 8 experts 02 Select Select top-k: E2 and E7 (k = 2) 03 Transform E2 E7 E₂(xₜ) E₇(xₜ) 04 Aggregate yₜ = 0.6 E₂(xₜ) + 0.4 E₇(xₜ) Conditional computation expands total capacity while each token uses only a few experts. Experts are not preset “math” or “writing” teams; routing and specialization emerge together during training. Teaching example: 2 of 8, with illustrative scores and weights. K3 uses 16 of 896 plus a shared path.
pⱼ = sⱼ / Σᵢ∈A sᵢ,yₜ = Σⱼ∈A pⱼ Eⱼ(xₜ)
Example: p₂ = 0.9 / 1.5 = 0.6; p₇ = 0.6 / 1.5 = 0.4
Why MoE deserves a closer look
Models benefit from greater parameter capacity, but computation is limited. MoE offers a tradeoff: maintain many feed-forward parameters while letting each token use only a small subset.
Inputs differ. Different tokens can benefit from different transforms, allowing the router to send them to different experts and letting specialization develop over training.
More experts make load balancing, communication across devices, and implementation more demanding. A FLOPs formula alone does not capture whether routing is balanced or kernels run efficiently.
An “expert” is a specialization learned during training. Each Eⱼ has its own weights and receives a different mix of routed inputs, gradually developing different processing preferences. The name does not require human-assigned categories such as “math expert” or “legal expert,” and an expert does not run its own conversation.
Which computation does sparsity save?
Suppose experts have similar computational cost, with N available and k activated per token. Total expert parameters grow with N, while expert computation per token mainly grows with k. Different tokens in a training batch can activate different experts, so most experts may still receive training within that batch.
| Question | Main determinants |
|---|---|
| How much expert parameter capacity can the model hold? | The combined weights of all N experts |
| How much expert computation does one token execute? | The selected k experts, plus the router, shared branch, and projections |
| How much storage or distributed capacity do the weights require? | All experts and the deployment scheme; a low activation fraction does not eliminate weight storage |
| How fast does it actually run? | Matrix shapes, batch size, device communication, load balance, and kernel efficiency |
Each MoE sublayer in K3 has 896 routed experts, with 16 activated per token, plus feed-forward capacity equivalent to two full-width shared experts used by every token. The LatentMoE chapter explains which vector width each path receives and how the shared branch connects.[2] [3]
AttnRes: revisiting information from earlier depths
Begin with residual accumulation, then examine content-dependent selection and the need for blocks.
A conventional residual connection is h_next = h + F(h). F(h) is the new information produced by the current sublayer, while h carries earlier information. Expanding the recurrence expresses the input at a given depth as the sum of the input embedding and earlier sublayer outputs. This direct additive path helps train deep networks, but it merges all sources into one state, making source selection less direct for later sublayers.[22] [25]
Fix token position t and compare depth paths: direct accumulation versus content-dependent reweighting. FIG 05 / Depth retrieval AttnRes: retrieving earlier representations at one position Fix token position t and compare depth paths: direct accumulation versus content-dependent reweighting. States Experts Depth Compute Token position t is fixed throughout Residuals: direct accumulation Input-dependent outputs; additive coefficient 1 AttnRes: content-based weights Weights change with the input content Earlier state h₀ Accumulated state h₁ Current state h₂ + Sublayer output f₀(h₀) × 1 + Sublayer output f₁(h₁) × 1 Depth Earlier block b₀ Earlier block b₁ Current partial sum b₂ Weight by content Σ αᵢ bᵢ Current sublayer input Read blocks at the same position AttnRes reads earlier depths of the same token, rather than other token positions. Residual coefficients are fixed; AttnRes computes content-dependent weights. K3 sums within blocks and weights across them. Sources: Attention Residuals, K3 report §2.2, and _apply_attn_res.
Attention Residuals (AttnRes) lets each sublayer read outputs from earlier depths and weight those sources individually. Selection occurs along depth. Conventional sequence attention reads candidates from different token positions; AttnRes reads representations left at different depths by the same token.[22] [25]
aᵢ,ₗ,ₜ = exp(eᵢ,ₗ,ₜ) / Σⱼ exp(eⱼ,ₗ,ₜ)
hₗ,ₜ = Σᵢ aᵢ,ₗ,ₜ uᵢ,ₜ
The same sublayer can assign different depth weights to different tokens because their uᵢ,ₜ differ. RMSNorm is used when computing matching scores, reducing the dominance of high-magnitude sources. The weighted aggregation uses the corresponding original source vectors.[25]
Why group layers into blocks?
Keeping every earlier sublayer output increases activation storage and communication across devices. Block AttnRes sums sublayer outputs within a depth block and saves the result as a block summary, while maintaining a partial sum for the current unfinished block. The next sublayer reads the input embedding, completed block summaries, and the current partial sum.[4] [25]
The diagram shows two completed blocks and one in progress. After a read, the new Attention or FFN output is added to the current partial sum. At a block boundary, that summary is saved and a new partial sum begins.[3] [25]
K3's 93 layers form seven complete 12-layer blocks and one final 9-layer block. Including the input embedding gives nine depth sources for the final aggregation. These summaries retain the token axis: each position is processed separately, rather than collapsing the entire sentence into one vector.[2] [4]
Inside an expert: GLU → SwiGLU → SiTU
Follow the two branches to make the idea of a gate concrete.
Inside an expert, two different linear projections of x produce equal-length vectors a = Wg x and b = Wu x. Think of b as the values to transmit and g(a) as coordinate-wise modulation of those values. After elementwise multiplication, the output matrix Wd restores the input width. Bias terms are omitted throughout this discussion.[17] [18]
All three rows share the same two-branch structure. GLU uses a sigmoid gate; SwiGLU introduces a within the gate; SiTU-GLU compresses the dynamic range of both the gate and value branches.[17] [18] [1]
GLU intermediate: z = σ(a) ⊙ b
SiLU(a) = a ⊙ σ(a)
SwiGLU intermediate: z = SiLU(a) ⊙ b
Complete FFN: E(x) = Wd z
The curves are calculated from the displayed formulas. SiLU grows approximately linearly for large positive inputs, while SiTU smoothly approaches 4. Both can produce negative gate values, so neither should be interpreted directly as a probability.[18] [1]
| Gate input a (value branch fixed at b=3) | GLU z | SwiGLU z | SiTU-GLU z |
|---|---|---|---|
| -2 | 0.358 | -0.715 | -0.658 |
| 0 | 1.500 | 0.000 | 0.000 |
| 2 | 2.642 | 5.285 | 4.861 |
| 8 | 2.999 | 23.992 | 11.509 |
For a=2 and b=3, GLU modulates the value by σ(2)≈.881. SwiGLU uses 2σ(2)≈1.762, so its output can exceed b. SwiGLU retains a larger positive dynamic range, but also allows large values in both branches to multiply at the same coordinate. This amplification path is what K3 aims to control.[18] [1]
SiTU-GLU bounds both branches before multiplication
SiTU(a) = 4 tanh(a/4) ⊙ σ(a)
z = SiTU(a) ⊙ [25 tanh(b/25)]
E(x) = Wd z
Bounding both branches constrains the intermediate product. Wd then recombines those coordinates, so the complete expert output also depends on the magnitude of its output weights.[1] [3]
To isolate function shape, set a=b>0. At a=40, the SwiGLU intermediate product is about 1600, versus 92.17 for SiTU-GLU. The vertical axis is logarithmic.[1]
Hard clipping has zero local derivative beyond its threshold. The real-valued derivative of softcap is sech²(v/β), which decreases smoothly. Floating-point implementations can still saturate numerically at extreme inputs. The authors report that softcap generally works better at the same bound; this is an observation within their experimental settings.[1]
LatentMoE: trading width for more experts
Expert count, active count, and expert width must be considered together.
A latent representation processes the input in a smaller feature space. LatentMoE shares a down-projection and an up-projection around its routed experts: it maps d-dimensional inputs to d/2, runs the experts there, and returns to d. If the expert intermediate width stays fixed, halving its input and output width halves the parameters in its three main matrices.[16]
The standard design is a budget baseline. The router selects experts from the original input; the projected features enter those experts; the shared branch stays full-width. The stable version inserts RMSNorm after routed aggregation and before the up-projection.[1] [3] [16]
Follow one token through the animation to connect the paths in the diagram, then calculate the parameter budget of that computation.
Watch the router read the full input, then follow its projected representation into the selected experts. Their weighted outputs pass through RMSNorm and the up-projection before being added to the full-width shared output. At the end, the network parameters stay fixed while the input changes, altering expert selection and the final result.
y = W↑ RMSNorm(u) + Shared(x)
With K3's width d=7168 and intermediate width m=3072, a full-width expert has 3dm main-matrix parameters, versus 3(d/2)m for a low-dimensional expert. Consequently, activating 8 of 448 full-width experts and 16 of 896 low-dimensional experts gives exactly the same active routed-expert matrix budget. The additional projections must still be counted separately.[2]
Bar lengths show the main-matrix parameters involved per token. The two added projections total about 51.4M, or 9.7% of the active routed-expert matrices. The shared-expert capacity, identical in both designs, is shown separately. Router, normalization, and communication costs are excluded.[2] [16]
What does the additional Norm control?
The low-dimensional routed path adds down- and up-projections around the expert's own matrices and gate. Their scales can compound. RMSNorm gives the up-projection an input with a controlled overall scale, reducing magnitude accumulation along this chain.[1]
The two-dimensional example sets γ=1 and ignores ε. It maps u and 3u to the same direction and magnitude. Actual training also involves learned gains, directions, gradients, and other factors.[1] [3]
The authors report more stable training and improvements on some downstream evaluations from placing Norm here. The balance between routed and shared outputs, normalization's nonlinearity, and optimization dynamics may all contribute. Observing a benefit and identifying its cause require different evidence.[1] [4]
Loss-Free and QB: balancing expert workloads
From a training feedback loop to thresholds derived from a distribution.
A router may favor a few experts, overloading them while leaving others with little training signal. A conventional remedy adds a balancing loss to the task loss: L = Ltask + λLbalance. Loss-Free uses a separate control path: adjust expert-selection biases b from observed loads. The task loss and model training remain in place.[19]
The upper path is the forward computation. Below it, task backpropagation and load-bias feedback are separate. Biases affect Top-k selection, while aggregation weights still normalize the selected experts' raw scores.[19] [3]
pⱼ = sⱼ / Σᵢ∈A sᵢ,j∈A
SignSGD-style feedback: bⱼ ← bⱼ + u · sign(target load − observed load)
“Loss-Free” means avoiding an additional balancing-loss term. Changing the routing changes which experts execute, and therefore changes task gradients too. Since inputs differ, balance is usually defined over a batch of tokens; each individual token need not use every expert equally.[19]
QB turns bias updates into a quantile problem
With N experts and k selections per token, each expert should receive approximately k/N of the tokens. QB first finds each token's current selection threshold. It then considers how far one expert's scores lie from those thresholds across the batch, and derives a new bias from that distribution. This removes the SignSGD step size u.[4]
The left side finds a token's cutoff along the expert axis. The right side finds an expert's quantile along the token axis. The example fixes current cutoffs and ignores tied scores; a real update may also change the next step's cutoffs.[4]
mᵢⱼ = sᵢⱼ − αᵢ
b̂ⱼ = −Q₁₋ₖ/ₙ(m[:,j])
b_new = b̂ − mean(b̂)
With current token thresholds fixed, the quantile gives the bias corresponding to the target selection share. Once all expert biases change, token thresholds may change as well, so adjustment continues across steps. Removing a learning-rate-like step size does not by itself guarantee perfect balance at every step or for every input distribution.
How is the global distribution aggregated?
Averaging local medians cannot recover the pooled median. Histograms retain counts over intervals; adding those counts reconstructs the pooled distribution's cumulative interval counts.[1] [4]
Appendix D of the K3 report specifies the procedure: build a histogram of rᵢⱼ=αᵢ−sᵢⱼ over [min(b)−1, max(b)+1], then read its k/N quantile. This range exploits sigmoid scores being in (0,1). Counts are merged across workers and gradient accumulation before linear interpolation within the selected bin.[4]
The random seed is 20260909. Across these 41 observations, the largest MaxVio difference between exact-quantile QB and the 1000-bin estimate is 0.02734. This is one teaching experiment under a fixed distribution; different score distributions, batches, or update rules may behave differently.
An independent analysis identifies another caveat: under a particular quantile convention, alternating updates can stop at a suboptimal fixed point. In an enumerable 3×3 example, the fixed-point dual objective is 5, while the best one-to-one assignment scores 4, leaving a gap. Fast adjustment in one simulation and a general convergence guarantee must therefore be assessed separately.[24]
Revisiting the evolution of expert designs
Across rows: improvements. Between rows: different problems. The rows do not form one causal chain. FIG 04 / Design connections Expert designs: three directions, three different problems Across rows: improvements. Between rows: different problems. The rows do not form one causal chain. States Experts Depth Compute 01 Expert structure Which parameters run? Common and conditional transforms Sparse experts Sparse MoE Activate a few on demand Shared experts Shared Expert Common transforms always run Low-dimensional experts LatentMoE Project down for expert work Stable latent experts Stable LatentMoE K3 combines these components Down Up Norm 02 Activations Inside one expert: how are magnitudes controlled? GLU Sigmoid on the gate branch SwiGLU Swish gate; unbounded product SiTU-GLU (K3) Smoothly bound both branches 03 Load balance Across routed experts: who is busy, who is idle? Auxiliary balancing loss Add an objective for balance Loss-Free Adjust loads with selection biases QB (K3) Update biases from score quantiles Shared experts compute common transforms; load balancing regulates routed-expert workloads. Sharing and balancing have different roles. K3 combines expert structure, numerical stability, and routing balance. Sources: DeepSeekMoE, LatentMoE, GLU Variants, Auxiliary-Loss-Free Load Balancing, K3 report §2.3.
Attention caches: shared heads and compressed features
Start with the roles of Q, K, and V, then compare two structural constraints in a common frame.
Attention uses three vectors with different roles: Query (Q) describes what to look for, Key (K) is matched against that query, and Value (V) supplies the information retrieved after matching. Each head uses its own query. It scores visible historical positions, converts those scores to normalized weights with softmax, and retrieves a weighted combination of V.[6]
oₜ = Σⱼ aₜⱼ vⱼ
During autoregressive generation, historical K/V can be cached. MHA keeps separate K/V for each head. GQA groups query heads so each group shares one K/V set. MQA shares one set across all query heads. All three retain individual historical positions; they primarily reduce the number of KV heads.[6] [7] [8]
All columns fix historical position j. Q denotes query heads; the lower objects are that position's cache. FIG 06 / Cached representations GQA / MQA / MLA: what is actually cached? All columns fix historical position j. Q denotes query heads; the lower objects are that position's cache. States Experts Depth Compute GQA Share K/V within groups Q1 Q2 Q3 Q4 MQA All heads share one K/V set Q1 Q2 Q3 Q4 MLA Cache a low-dimensional latent Q1 Q2 Q3 Q4 Kⱼ / Vⱼ Kⱼ / Vⱼ Position j: 2 K/V groups Share across heads One shared Kⱼ / Vⱼ Position j: 1 K/V group Maximum head sharing Low-dimensional latent cⱼ K/V K/V K/V K/V Position j: compressed features cⱼ The upper expansion denotes projection Compress feature dimensions GQA/MQA share KV heads; MLA changes the representation cached at each historical position. These are compressed-cache representations, not a requirement on a backend's physical storage format. Sources: MQA, GQA, DeepSeek-V2 §2.1, K3 report §2.1.2. Four query heads are used for a common comparison.
MLA factorization: K₁ = U₁ᴷ c, K₂ = U₂ᴷ c; likewise for V
A minimal example: let c=(1,2), and let two projections read its first and second coordinates. They produce K₁=1 and K₂=2. Two heads in the same GQA group instead read the same K. Thus, projecting a common source differently and directly sharing one K/V set impose different expressive constraints.
How does MLA retain multiple heads with a small cache?
Each K3 MLA layer has 96 query heads. Its compressed representation contains a 512-dimensional latent c and a 64-dimensional shared-key path r. Each head can project c into a 128-dimensional content Key and a 128-dimensional Value, then append r to the Key to obtain 192 dimensions.[2] [3]
Both sides compute the same function. The left expands each head's K/V first. The right moves the content-Key up-projection to the Query side and aggregates Values in latent space. Both preserve the original scaling, causal mask, and softmax.[3] [9]
Σⱼ aⱼ(Uⱽcⱼ) = Uⱽ(Σⱼ aⱼcⱼ)
K3 adds an elementwise gate g(x) after concatenating the head outputs and before the final projection Wo. Uⱽ can be applied after aggregation. Because g(x) varies with the input and modulates individual coordinates, Uⱽ generally cannot be moved across it and statically fused into Wo.[3]
Combining KDA and MLA
Two kinds of memory alternate across layers: a fixed-size state and per-position records.
KDA writes history into a continuously updated associative-memory matrix S. Key determines the write direction, Value supplies the information to remember, and Query reads the current state. A delta update first retrieves the old prediction for the current Key, then corrects memory by the difference between the new Value and that prediction. KDA also controls retention separately across channels.[10] [11] [12] [13]
The diagram uses S∈ℝᵈᵏˣᵈᵛ. Channel-wise retention comes first, then a delta update at the current Key, followed by a Query read. The actual KDA module also includes short convolution, normalization, and output gating.[13] [3]
v̂ₜ = Ŝₜᵀ kₜ
Sₜ = Ŝₜ + βₜ kₜ(vₜ − v̂ₜ)ᵀ
õₜ = Sₜᵀqₜ
An additive memory simply adds the new Value to the old one. A delta update attempts to correct the existing association. If a unit Key previously maps to 1 and the new Value is 3, the teaching case with no forgetting and a full-strength write gives 1+(3−1)=3. This illustrates how the rule can overwrite an existing association.
The hybrid operates across network depth
Alternate two sequence mixers across depth: recurrent history compression and content reads by position. FIG 07 / Hybrid memory KDA + MLA: summary states and per-position reads Alternate two sequence mixers across depth: recurrent history compression and content reads by position. States Experts Depth Compute Repeating pattern: 3 × KDA + 1 × MLA Sequence-mixing sublayers only Sublayer 1 KDA Sublayer 2 KDA Sublayer 3 KDA Sublayer 4 Gated MLA A repeating pattern, not K3's total layer count. A complete layer also has FFN and AttnRes. KDA: update a fixed-size state x₁ x₂ … xₜ S₁ S₂ Sₜ₋₁ Sₜ Write new information into a summary; decode-state shape stays fixed. Compressed state cannot preserve every positional detail without loss. MLA: a low-dimensional cache per position c₁ c₂ c₃ … cₜ₋₁ xₜ Keep per-position reads; the cache grows with sequence length. KDA compresses history into state; MLA retains a per-position read path. The left shows sequence mixers only. Complete layers also contain FFNs, with AttnRes providing depth connections. Configuration: 69 KDA and 24 MLA layers, 93 in total. The final layer also uses MLA.
One layer's output representations become the information available to the next layer. KDA's order information reaches later MLA layers through representations and depth connections. MLA reads the representations at historical positions in its own layer. Each KDA layer maintains its own recurrent matrix.[3] [13]
Kimi Linear already used a repeating pattern of three KDA layers followed by one MLA layer. K3 places MLA at layers 4, 8, …, 92, and also at the final layer 93, giving 69 KDA layers and 24 MLA layers. The first layer has a dense feed-forward network; the others use MoE.[2] [13]
Each vertical column is a complete Transformer layer. The upper row identifies its sequence mixer and the lower row its feed-forward network. Brackets show AttnRes blocks of 12 complete layers, with 9 layers in the final block.[2] [3]
A fixed state still has a cost
M_hybrid(T) = 24 × 576 × T × 2 + 69 × 96 × 128² × b bytes
Only the main attention state for a single sequence is compared, assuming compressed MLA caching. The two hybrid curves show the effect of KDA state precision. Logarithmic axes support comparison across lengths. These are not measured memory footprints from a specific backend.[2]
The formulas give crossover points of approximately 2731 tokens for b=2 and 5461 tokens for b=4. The comparison balances cache growth in the 69 replaced MLA layers against the added fixed KDA states. At long contexts, the hybrid's cache-growth slope is determined by its 24 MLA layers.
Two historical attention directions meet here
The upper path retains historical positions while exploring KV sharing and low-dimensional compression. The lower path studies how to update a fixed-size state. Kimi Linear combines KDA and MLA across layers, and K3 retains and adjusts that structure. Horizontal spacing is for layout only.[6] [7] [8] [9] [10] [11] [12] [13] [1]
RoPE and NoPE: where order information comes from
First calculate how rotation changes a score, then examine K3's specific choice.
In attention, qₘᵀkₙ measures content matching. Rotary Position Embedding (RoPE) explicitly adds relative position by rotating two-dimensional coordinate pairs in Query and Key. Later positions rotate farther at the corresponding frequency, with different coordinate pairs using different frequencies.[20]
RoPE explicitly adds relative displacement. K3's NoPE removes explicit rotation from the MLA branch. FIG 08 / Position and scope RoPE and NoPE: how position enters the computation RoPE explicitly adds relative displacement. K3's NoPE removes explicit rotation from the MLA branch. States Experts Depth Compute RoPE: position rotates vectors A single two-dimensional subspace Position m Position n qₘ q′ₘ kₙ k′ₙ Angles and vector lengths are illustrative. q′ₘ = Rₘ qₘ k′ₙ = Rₙ kₙ q′ₘᵀ k′ₙ = qₘᵀ Rₙ₋ₘ kₙ Rotation depends on relative displacement n − m; content vectors q and k also determine the score. K3: remove explicit RoPE from MLA x₁ x₂ x₃ xₜ S₁ S₂ S₃ Sₜ KDA: update in sequence order Token states from an order-sensitive path These states pass through layers for MLA to read. MLA: content reads over visible positions Causal visibility: self and past only. MLA does not explicitly rotate Q/K by position. K3 removes RoPE inside MLA in a hybrid design; this is not a universal result for pure attention. It does not guarantee cost-free removal in any model, or exact preservation of all relative-position information. Sources: RoFormer/RoPE, K3 report §2.1.2, mla_use_nope=true, and the KDA implementation.
This example uses one two-dimensional coordinate pair. To isolate position, all three content vectors are fixed at (1,0). The angle θ=30° is illustrative; actual RoPE combines multiple rotation frequencies.[20]
NoPE retains content attention and causal visibility
NoPE:ℓₘₙ = [qₘᵀ kₙ] / √dₖ + Mₘₙ
aₘₙ = softmaxₙ(ℓₘₙ)
With content fixed, NoPE gives the same dot product at every visible position, while RoPE varies with relative position. In a real model, q/k have already been processed in context, so NoPE scores can also reflect order encoded by earlier layers.[20]
The dot product in this two-dimensional example decreases between 0° and 60°. This does not imply that RoPE scores decay monotonically for every direction and distance. Rotation is periodic, and actual scores also depend on content vectors and the combination of frequencies.
Why can K3's MLA use NoPE?
KDA layers are interleaved before MLA. Recurrence and short convolution depend on input order, so MLA's inputs can already contain order information even when MLA no longer rotates Q/K. The causal mask also preserves the structure of the visible prefix. A two-number recurrence makes this easier to see.[3] [13]
This scalar recurrence shows only how order can enter a hidden representation. Actual KDA uses a matrix state with input-dependent forgetting, writing, and output gates. The example does not quantify K3's ability to identify positions.
The source reports that adding RoPE back to MLA in K3's KDA+MLA hybrid showed no clear benefit, while removing RoPE from the all-MLA K2 substantially degraded performance. The effect of removing position rotation depends on the overall architecture and experimental conditions. A causal mask and order-sensitive recurrence alone do not guarantee exact-position reasoning or successful extrapolation to arbitrary lengths.[1]
Under what conditions are RoPE, PaTH, and KDA related?
KDA state transition: Aₜ = (I − βₜ kₜkₜᵀ) Diag(αₜ)
Powers of a fixed orthogonal transform can define generalized position transforms. PaTH explores input-dependent Householder-style path transforms. The delta rule also contains I−βkkᵀ, which can be related to an orthogonal reflection under specific normalization and coefficient conditions. KDA adds channel-wise decay. These connections establish relationships between mechanisms; their effects in complete models still require experimental evidence.[1] [21]
What does retaining the 64-dimensional path mean?
K3's 512-dimensional latent expands into per-head content Keys and Values. The 64-dimensional path is appended directly to every head's Key. It continues contributing to content scores after RoPE is removed.[1] [3]
Key-path score in K3 NoPE: (qₘʳ)ᵀ rₙ
Complete K3 logit per head: [(qₘᶜ)ᵀUᴷcₙ + (qₘʳ)ᵀrₙ] / √192 + Mₘₙ
The authors retain the path for compatibility with existing MLA infrastructure and to avoid the additional computation of expanding all 576 latent dimensions into each head's K/V. Long-context capability also depends on training data and schedule: the report describes successive length stages of 8K, 64K, 256K, and 1M.[1] [4]
Separating cache size, computation, and task time
The same design can save storage while increasing a particular computation.
Prefill processes new input tokens in a batch; Decode generates output tokens step by step. Expanded MLA uses relatively small per-head dimensions, while compressed-cache MLA projects Query into the larger latent space. These computation orders have different core costs.[1] [9]
For each valid query–key position pair, the main QK and weighted-V multiply-accumulate coefficient is H(dqk+dv). Holding H=96 isolates the listed head dimensions. Projections, softmax, and other modules are omitted.[1] [2]
The chart counts cached scalars per historical position per attention layer. GQA8 means 8 KV heads. Shared K=V stores one 512-dimensional vector. MQA 256/256 simplifies the attention-core dimensions used in the source's discussion of MFA.[1] [2]
Compressed MLA uses 576 dimensions for QK and 512 for aggregated latent Values, giving 96×(576+512)=104448. This counts only the main operations that depend on historical positions; output expansion and other operations must be added separately.[1] [2]
One MAC is a multiplication followed by accumulation. The three charts use different units: the MAC coefficient describes part of the computation, while scalar count describes cache capacity. Generation time also depends on memory reads, matrix utilization, batch size, communication, and scheduling. A 4.25× core MAC coefficient does not directly imply 4.25× latency.
Speculative decoding changes the value of extra computation
Speculative decoding uses a smaller draft model to propose several tokens, then asks the target model to verify them in a batch. Its gains depend on verification efficiency and acceptance rate, offset by draft computation and state-maintenance costs. MLA's larger latent-space computation may hit a compute bottleneck sooner in some settings. KDA's recurrent state also requires rollback handling after rejected proposals.[1] [4]
K3's training and release configurations describe different stages. Pretraining includes one MTP layer, which is later fine-tuned into an EAGLE-3-style draft model. The released main-model configuration has num_nextn_predict_layers=0. The first statement describes training history; the second describes the released main model.[2] [4]
Another approach: DeepSeek-V4 compression and sparse access
K3 uses many fixed-state layers to reduce per-position caching. V4 combines local windows, compressed historical entries, and sparse access. CSA compresses before selecting entries, while HCA applies stronger compression. Their access patterns must be considered separately.[1] [5]
| Architecture field | DeepSeek-V4-Flash | DeepSeek-V4-Pro |
|---|---|---|
| Complete Transformer layers | 43 | 61 |
| Query heads | 64 | 128 |
| Shared KV dimensions | 512 | 512 |
| Total / active parameters per token | 284B / 13B | 1.6T / 49B |
These fields come from the V4 technical report. Its shared K=V structure is related to the latent-space view of MLA decoding, while compression and sparsity introduce further tradeoffs. Fixed matrices can cause memory interference, compressed entries can merge local details, and sparse reads can miss relevant records. Choosing between these combinations requires task evaluations under comparable conditions.[5] [1]
Connecting the mechanisms to K3's design choices
We can now connect the three directions from the opening diagram and assess the architectural tradeoffs.
A supporting training direction: Muon and Per-Head Muon
The preceding modules define the forward computation; the optimizer defines how training updates the weights. Muon specially transforms matrix-shaped update directions, and Moonlight extends it to large-model training. The per-head form processes relevant attention updates in head-wise blocks. The source reports no clear performance difference from this change, presenting it mainly as a way to avoid unnecessary coupling.[1] [23]
How can more experts have a similar cost?
Total expert count N, active count k, and expert width jointly determine the budget. K3 halves the routed experts' input/output width while doubling both N and k, preserving the active budget of their three main matrices. Down- and up-projections, routing, and communication still add costs.
Why can decoding remain costly with MLA's small cache?
The shared latent saves storage, but each Query head processes longer vectors when scoring and aggregating in latent space. Actual speed depends on the balance between computation and memory access; speculative decoding changes how both are utilized.
What supports position capability without RoPE?
KDA recurrence, short convolution, and causal visibility provide paths for order information. Later NoPE MLA layers can read representations that already contain context. Whether RoPE helps, and how well position tasks or length extrapolation work, requires experiments on the relevant architecture.
Looking back across the design directions: MoE enables conditional feed-forward computation; MLA separates multi-head representation from cache width; KDA provides fixed-state memory; AttnRes adds depth-wise selection. Stable LatentMoE, output gates, position-related branches, and optimizer changes help these components work together in one system.[1]
where the data comes from, what computation it undergoes, which resource is saved, and what cost is added.
References and calculation notes
Sources, teaching derivations, and experimental observations are cited according to their roles.
This article follows Jianlin Su's architectural discussion of K3. Model dimensions and switches come from the released configuration, and computation paths from the public reference code. Claims about design effects retain the scope of the corresponding papers or author observations. Small numerical examples, function curves, and capacity models explicitly state their teaching assumptions.
Architectural reasoning, tradeoffs, and experimental observations.
text_config: complete layers, experts, head counts, dimensions, and switches; commit c5d1dd4 as listed on the page.
SituAndMul, KimiSparseMoeBlock, KimiMLAAttention, KimiDecoderLayer, and _apply_attn_res.
Version dated 2026-08-07; architecture, MTP draft layer, QB derivation, and histogram estimation.
Shared KV, compressed sparse attention, and Flash/Pro configurations.
Transformer architecture, multi-head attention, and feed-forward networks.
Multi-Query Attention with shared K/V.
Grouped-Query Attention with K/V sharing within groups.
Associative memory and delta updates.
Definitions and experiments for SwiGLU and other GLU variants.
Independent analysis: fixed-point counterexamples for quantile updates and extensions to load constraints.
Block AttnRes pseudocode in the README. K3's block-size convention follows its own configuration and implementation.
Scope of the teaching calculations
Save the calculation code (Python / NumPy). Run python K3_Calculation_Examples.py --output results.json to calculate expert parameter counts and state capacity, and check the MLA, KDA, RoPE, and QB teaching examples.
Expert parameters count the three main matrices. Attention cost charts include only core multiply-accumulates, and the memory model includes only the listed attention states. These reveal structural trends but cannot replace measured throughput or full memory tests. The QB curves use a fixed distribution with resampled batches. The accompanying code reproduces these examples and matrix-equivalence checks.
This article does not report retraining K3, loading the complete model, or reproducing its ablation experiments. Experimental claims retain the scope specified by their sources.