Jonas Hansen

How modern language models work

The whole article as a progression of ideas, one line each. Every step links to its full section.

From text to vectors

  1. Text A language model never reads letters directly: before it can do anything, text has to become a sequence of integers.
  2. Tokenization Text is split into units from a fixed vocabulary: whole words, pieces of words, or raw bytes.
  3. Token IDs Each token is replaced by its position in the vocabulary. The number is an index, and carries no meaning of its own.
  4. Embeddings Token IDs select rows of a learned table of vectors: the model's starting point for each token, not its final word on meaning.
  5. Positional information Attention on its own ignores word order, so position has to be injected. Most current models do it by rotating queries and keys (RoPE).

Inside the transformer

  1. Context-dependent representations As vectors pass through the network, each token's representation is rewritten using the tokens around it.
  2. Transformer layers A transformer is a stack of blocks. Each one lets tokens exchange information (attention), then transforms every token on its own (an MLP).
  3. Self-attention Each token's representation gathers information from other positions in the context, weighted by how relevant the model computes them to be.
  4. Queries, keys and values Each token is projected into a query (what it looks for), a key (what it can be matched on) and a value (what it passes along).
  5. Dense attention In full (dense) attention every token can attend to every earlier token, so the work grows with the square of the sequence length.
  6. Sparse and sliding-window attention Some attention layers restrict each token to a window or pattern of positions, trading direct reach for lower cost. It all happens inside the model.
  7. Context windows The context window is the most tokens a model can process at once: prompt, documents, tool results and its own output together.
  8. KV cache During generation, keys and values already computed for earlier tokens are kept and reused, so each new token only computes its own.

Choosing the next token

  1. Logits The final representation at the last position is turned into one score per vocabulary entry: the logits.
  2. Probabilities and softmax Softmax turns the logits into a probability distribution over the entire vocabulary.
  3. Decoding and sampling Generation is a loop: choose one token from the distribution, append it to the input, run the model again.
  4. Temperature, top-k and top-p Temperature reshapes the distribution; top-k and top-p cut off its unlikely tail before a token is drawn.
  5. What happens to the tokens not chosen Next-token candidates that were not chosen are normally discarded. The model does not keep them as alternative lines of reasoning.
  6. Multiple samples, search and verification Systems can explore alternatives deliberately: generate several candidates, search over partial solutions, and check the results. It is an outer loop around generation.

Bringing in outside information

  1. Retrieval-augmented generation RAG searches an external collection at request time and places the most relevant passages in the context before the model answers.
  2. Retrieval versus attention Attention mixes information among tokens already in the context. Retrieval decides which outside information gets into the context at all.
  3. External memory Anything a system should remember beyond one context window has to be stored outside the model and read back in as tokens.
  4. Tool use A model can emit a structured request to call a tool. Software outside the model runs it and feeds the result back into the context.
  5. Model Context Protocol (MCP) MCP is an open protocol for exposing tools, data and prompts to AI applications. It standardises the connection, not the retrieval.
  6. Codebase-aware retrieval Coding agents find the relevant parts of a repository with search tools, such as text search, symbol indexes, embeddings or code graphs, and read them into context.

Learning, and what the words mean

  1. Inference versus training Training changes a model's weights. Inference, meaning ordinary use, runs fixed weights and changes nothing in the model.
  2. Self-training and synthetic data Model outputs that have been filtered, checked or ranked can become training data for a later model: a learning loop outside generation.
  3. Recursive self-improvement The idea that an AI system could improve its own capabilities, including its ability to improve itself. A research concept, not a current mechanism.
  4. AGI, and why definitions vary Artificial general intelligence has no single agreed definition, so any claim about it depends on which definition is being used.