Column

Structured long-form essays, mechanism breakdowns, and interactive notes.

  1. K3 Architecture Illustrated
    K3 Architecture Illustrated

    Understand K3's experts, attention, and depth-wise information flow through the computation of a single token: illustrated chapters, derivations, numerical examples, and resource tradeoffs.

    Read
  2. The article’s opening screen: the headline “How does one token’s cache become a shared memory?” beside a diagram of input history, state lookup and a shared global KV
    Xiao Wang Asks, Da Wang Answers: From KV Cache to YOCO and Shared Memory

    22 questions and answers between Xiao Wang and Da Wang, with diagrams: the KV cache, chained hashes, Claude’s cache breakpoints, DeepSeek’s persistence, YOCO’s shared memory, an 890-byte conditional ledger, approximate restoration, sparse retrieval and Engram. 17 mechanism diagrams and 4 offline interactive labs.

    Read
  3. Technical flow diagram showing a question moving through tokens, context, tool calls, sampling, and output generation in an LLM
    A Question's Journey - LLM from Input to Output

    A full mechanism timeline from one prompt through context, tokenization, Transformer internals, generation, tool calls, sampling, latency, and reasoning budget.

    Read