learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

How an LLM Writes Text

What really happens between a prompt and an answer: tokens and the tokenizer, one score per vocabulary entry, softmax, the generation loop that runs the whole model once per token, greedy decoding, temperature, top-k and top-p, seeds, and the three ways generation stops.

An interactive AI Infrastructure lesson: 20 steps, about 30 minutes, on a live simulation in your browser.

Every LLM you will run in production, from Llama 3.1 8B to the largest frontier models, does one thing: given some text, it scores every possible next token, and something picks one. Then it does it again. That loop is what you are paying GPUs for.

To watch it, the stage runs a model small enough to read. It was trained on 45 sentences from an infrastructure team's chat, and it predicts from the last two tokens only. A real model reads the whole conversation through billions of learned weights, but the loop around it is exactly the same.

What you will learn

  1. Text in, one token out

    • A language model you can read: An LLM never writes a sentence at once: it predicts one next token, appends it, and repeats.
  2. Tokens, not words

    • Text becomes token ids: A tokenizer turns text into ids from a fixed vocabulary; the model only ever sees those ids.
    • Capitals, and a character it never saw: A tokenizer has a piece for every character in its base alphabet and learned pieces for common words; anything outside the vocabulary falls back to raw bytes, so it always works but costs more tokens.
    • Tokens are the unit of everything: Context limits, prices, throughput and KV cache memory are all counted in tokens: estimate with ¾ of a word per token, measure with the real tokenizer.
    • Drill: words to tokens: Tokens ≈ words ÷ 0.75 for English text: 3,000 words is about 4,000 tokens.
  3. Scores and probabilities

    • One score for every token: A forward pass returns one logit per vocabulary token; softmax, p_i = e^(z_i) / (Σ_j e^(z_j)), turns them into probabilities that add up to 1.
    • How sure is the model?: The model returns a distribution, not an answer; at most positions several tokens are plausible, and the decoding strategy decides which one you see.
    • Drill: softmax by hand: Softmax: p_i = e^(z_i) / (Σ_j e^(z_j)). Adding the same number to every logit changes nothing; only the differences matter.
  4. Again and again

    • Append it and run again: Autoregressive: each new token is predicted from the prompt plus every token generated so far, then becomes part of the input.
    • How many passes for one answer?: Output length is sequential work: N tokens means N forward passes one after another, which is why long answers take long no matter how many GPUs you have.
    • Break it: overrule one token: Tokens are never revised: one early choice conditions everything after it, which is how a single unlucky token can derail a long answer.
  5. Choosing the next token

    • Greedy: always the top token: Temperature 0 is greedy decoding: no randomness, the same output for the same input, every time.
    • Sampling: draw from the bars: Sampling draws the next token in proportion to its probability; the seed makes the draws repeatable.
    • Turn the temperature up: Temperature rescales logits: under 1 sharper, over 1 flatter; set it high enough and any model drifts into nonsense, because the tail of unlikely tokens takes over.
    • Top-p: cut the tail first: Top-p keeps the smallest set of likely tokens that reaches p and redraws from them; it removes the long tail that causes nonsense without making the output deterministic.
  6. Stopping

    • Break it: ignore the end token: Generation ends on an end-of-sequence token, a stop string or max_tokens; take away the first and greedy decoding loops until the budget runs out.
    • Break it: overflow the window: Prompt tokens plus max_tokens must fit the context window, or the request is rejected before any work is done.
    • What two tokens cannot do: The generation loop is the same for every LLM; what makes a model good is how much of the context its scores can use, and transformers use all of it through attention.
  7. Recap & playground

    • Cheat sheet
    • Playground