learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Inside the Transformer

What sits between token ids and next-token scores: embeddings, attention as queries, keys and values, the √d scaling and softmax, position information, the causal mask and why it makes a KV cache possible, many heads, the MLP, the output head, and where a real model's parameters, FLOPs and memory go, including mixture of experts.

An interactive AI Infrastructure lesson: 19 steps, about 30 minutes, on a live simulation in your browser.

"How an LLM Writes Text" ended on this: the tiny model continues "the engineer" with "checked the gpu is waiting for data." By the time it chose "is", it could see only "the gpu". The engineer was gone.

Any model that predicts from a fixed, short window fails like this. Real text needs the subject from ten words ago, the variable from forty lines up, the instruction at the top of a 100,000-token prompt.

What you will learn

  1. Two tokens are not enough

    • The model that forgot the subject: The generation loop is the same for every LLM; the transformer is the box that turns the whole context so far into next-token scores.
    • One forward pass, bottom to top: A transformer: embeddings → N identical layers (attention, then MLP, each added back to a residual stream) → final norm → output head → one logit per vocabulary token.
  2. Ids become vectors

    • The embedding table: An embedding is a learned vector per token id, looked up from a vocabulary × hidden table; the output head maps the final vector back to one score per vocabulary entry.
  3. Attention

    • Every token asks a question: Attention: each token computes a weighted mix of earlier tokens' information, with the weights chosen by how relevant each one looks to it.
    • Queries, keys and values: Attention(Q, K, V) = softmax(QKᵀ / √d)V: queries meet keys to make weights, and the weights mix the values.
    • Which noun does "it" mean?: Attention weights are driven by content: a token five or five thousand positions back is equally reachable.
    • Drill: one attention score: An attention score is (q · k) / √d; softmax over a row turns the scores into weights.
  4. Order and the mask

    • A head that tracks position: Attention by itself ignores word order; position embeddings (RoPE in Llama) rotate queries and keys so scores can depend on relative position.
    • Break it: take positions away: Without position information a transformer sees a bag of tokens; order exists only because positions are encoded into queries and keys.
    • What changes when a token arrives: With the causal mask, earlier tokens never see later ones, so their keys and values never change after they are computed: the reason a KV cache is possible.
    • Break it: let tokens see the future: A decoder without the causal mask could cheat in training and could not cache in inference; the mask is what makes generation both learnable and cheap.
  5. Heads, MLP and layers

    • Many heads, one layer: Multi-head attention runs many small heads in parallel on slices of the vector and recombines them; grouped-query attention lets several query heads share one key/value head.
    • The MLP: each token on its own: Attention mixes information across tokens; the MLP processes each token independently and holds most of the parameters.
    • Where 70 billion parameters live: In large dense transformers about 80% of parameters are MLP weights, about 17% attention, and the embeddings become negligible.
  6. From diagram to bill

    • Two FLOPs per parameter per token: For a model with N parameters, a forward pass costs ≈ 2N FLOPs per token; training adds the backward pass, ≈ 6N per token.
    • Drill: FLOPs for one token: 70B-class model: about 141 GFLOPs per token; at 1,000 tokens per second that is 141 TFLOPS of useful work.
    • Mixture of experts: big but sparse: Mixture of experts: total parameters set the memory, active parameters set the FLOPs; a router picks a few experts per token per layer.
  7. Recap & playground

    • Cheat sheet
    • Playground