Jonas Hansen

How modern language models work

From text to tokens, attention, sampling, retrieval and tools: how a large language model turns a prompt into an answer, and what it does not do along the way.

· 29 sections

From text to vectors

A computer stores text as bytes: Unicode characters encoded, usually, as UTF-8. A neural network does not compute with bytes or letters. It computes with numbers arranged in vectors and matrices, and the only thing it can do with an input is multiply and add it.

So between the prompt you type and the first calculation inside the model there is a fixed, mechanical pipeline. The text is normalised, split into tokens, and each token is replaced by an integer ID. Those integers are the model’s entire view of your text.

Nothing in that pipeline is learned while you use the model, and nothing in it understands anything. It is closer to a dictionary lookup than to reading.

This has consequences that surprise people. A model that receives straw and berry as two IDs has no direct view of the letters inside them, so “how many r’s are in strawberry” asks about something the model only knows indirectly, from patterns it absorbed during training.

Often misread as:

  • “A language model reads the letters of a prompt the way a person does.”

A tokenizer splits text into units drawn from a fixed vocabulary, typically tens of thousands to a few hundred thousand entries. Modern tokenizers work on subwords: common words get a single token, while rare words are built from several pieces. Byte pair encoding (BPE), WordPiece and Unigram are the common algorithms; byte-level variants fall back to individual bytes, so any input at all can be encoded.

The vocabulary is chosen before the model is trained, from statistics of a large text sample, and it never changes afterwards. The model learns to work with whatever pieces the tokenizer produces.

This is the actual output of the WordPiece tokenizer behind this article’s Ask view. ## marks a piece that continues the previous one:

"Why doesn't a larger context window replace RAG?"
why | doesn | ' | t | a | larger | context | window | replace | rag | ?

"transformers hallucinate unfathomably"
transformers | hall | ##uc | ##inate | un | ##fat | ##hom | ##ably

Tokens are not words. Punctuation and word fragments are tokens too, and a different model’s tokenizer would split the same sentence differently, which is why token counts and per-token prices are only comparable within one model family.

Often misread as:

  • “One token is one word.”

Once text is split, each token is looked up in the vocabulary and replaced by its index. The sequence of these integers is what the model actually receives.

The numbers are labels, not quantities. Token 2,000 is no more similar to token 2,001 than to token 9. Nothing in the model ever does arithmetic on the IDs themselves; their only use is to select a row from the embedding table.

Vocabularies also contain special tokens that never appear as ordinary text: markers for the start and end of a sequence, separators, and in chat models the markers that delimit system, user and assistant turns. A model stops generating when it produces its end-of-sequence token, because it was trained to emit one where text ends.

Often misread as:

  • “Tokens with nearby IDs have similar meanings.”

The model’s first layer is a table with one row per vocabulary entry. Each row is a vector of a few hundred to several thousand numbers, and a token ID simply selects its row. The values in the table are learned during training, along with everything else.

Training tends to place tokens that are used in similar ways near each other, which is why embeddings are often described as encoding meaning. For this table, that description goes too far. The input embedding for bank is the same row whether the sentence is about a river or an account: it is a context-free starting point. What a token means in a particular sentence is only worked out as its vector is transformed by the layers above, using the tokens around it. That is the subject of contextual representations.

The word embedding is also used for something related but separate: an embedding model that turns a whole passage into one vector, so passages can be compared by similarity. That is how retrieval usually finds relevant text.

The Ask view of this article uses the simplest kind: a static embedding table, where a passage’s vector is the average of its tokens’ rows. No context and no layers, so it runs in a browser without a machine-learning runtime. It is also why Ask combines it with plain keyword matching: a static table cannot tell which sense of a word you meant.

Often misread as:

  • “An embedding is a fixed definition of what a word means.”

Self-attention compares tokens with each other, but nothing in that comparison depends on where they are. Without extra information, “dog bites man” and “man bites dog” would give the model the same tokens in a different order, and it would have no way to tell them apart.

Early transformers fixed this by adding a position vector to each token’s embedding: fixed sinusoidal patterns in the original Transformer, a learned table of positions in BERT and GPT-2.

Most current large models use rotary position embeddings (RoPE) instead. Nothing is added to the embedding. Inside every attention layer, query and key vectors are rotated by an angle that depends on their position, so the score between two tokens ends up depending on how far apart they are. Other schemes, such as ALiBi, add a distance-based penalty to attention scores directly.

The choice matters for context length: a model handles positions well mostly within the range it was trained on, and many long-context techniques work by adjusting how RoPE behaves beyond that range.

Often misread as:

  • “Every model adds a position vector to each token embedding.”

Inside the transformer

The vector a token starts with is the same in every sentence. The vector it has after the first layer is not. Each layer mixes in information from other positions and transforms the result, so by the middle of the network the vector at the position of bank is different in “river bank” and “bank account”. These intermediate vectors are called hidden states, and together they form what is often called the residual stream.

This is where anything resembling meaning in context lives. Research that probes these vectors finds they encode things like grammatical role, which entity a pronoun refers to, and facts associated with a name. None of that was looked up; it was computed from the input by fixed weights.

Contextual representations are not stored anywhere permanently. They are computed fresh for each input and discarded when the request ends. The only part of them kept around during a single generation is the attention state in the KV cache.

At the top of the network, the representation at the last position is what the model uses to score possible next tokens: the logits.

Often misread as:

  • “The meaning of a word in context is looked up from a stored table.”

A large language model is a stack of dozens of transformer blocks, often somewhere between 30 and 100 or more. Each block reads the residual stream, computes something, and adds its result back, so information accumulates rather than being replaced.

Every block has two parts:

  • Attention moves information between positions. It is the only place where one token’s representation can be influenced by another’s. See self-attention.
  • A feed-forward network (MLP) then transforms each position separately, with the same weights at every position. Much of what a model has absorbed about the world appears to be stored in these weights.

Normalisation steps keep the numbers in a stable range between the two.

The weights of each block are fixed after training and shared across positions, which is why one model can process a prompt of five tokens or five thousand: the same computation is applied at every position. Chat models are decoder-only transformers with a causal mask, meaning each position can only attend to positions before it.

Many recent large models replace the single MLP with a mixture of experts: several MLPs, of which a small router picks a few for each token. The model has more parameters in total, but each token only uses a fraction of them.

Often misread as:

  • “Each layer of the model handles a different word of the input.”

Attention is how a transformer moves information between positions. For each position, it computes a weight for every position it is allowed to look at, then takes a weighted average of information from those positions. In a decoder-only model the allowed positions are the current one and everything before it.

The weights are computed fresh for every input. That is the important difference from the rest of the network: the MLP weights are fixed numbers, but how much “it” draws from “the cat” versus “the mat” is decided at run time, from the content itself. How those weights are computed is the subject of queries, keys and values.

Each layer runs several attention heads in parallel, each with its own learned projections, so different heads can track different relations: the previous token, the subject of the sentence, a matching bracket.

Attention only ever operates over what is in the context window. It has no access to documents, databases or earlier conversations unless something has put them there as tokens. That is a job for retrieval, which is a different mechanism.

Often misread as:

  • “Attention lets the model look things up beyond what is in its context.”

Inside an attention head, each token’s current vector is multiplied by three learned matrices, producing three new vectors:

  • a query: what this position is looking for;
  • a key: what this position can be matched on;
  • a value: what this position contributes if it is attended to.

The score between position i and position j is the dot product of i‘s query with j’s key, divided by the square root of the key length to keep the numbers in a stable range. A softmax turns each row of scores into weights that sum to one, and the output for i is the weighted sum of the values:

Attention(Q,K,V)=softmax(QKTdk)V

The database vocabulary is only an analogy. Every key matches every query to some degree, and the output is a blend, not a single selected record.

Keys and values depend only on a token and the tokens before it, never on later ones, so once computed they can be reused. That is what the KV cache stores. Grouped-query attention lets several query heads share one set of keys and values, which shrinks that cache considerably.

Often misread as:

  • “Attention selects one matching token, like a database lookup.”

In the standard design, every position attends to every position before it. For a sequence of n tokens that is roughly n2/2 query-key scores per head, per layer. Double the length of the input and the attention work roughly quadruples, while the rest of the network’s work only doubles.

For short prompts this hardly matters. For hundreds of thousands of tokens it dominates, and it is one of the reasons context windows have limits at all.

A lot of engineering goes into making dense attention cheaper without changing its result. FlashAttention, for example, computes exactly the same output, but in tiles that stay in the GPU’s fast on-chip memory instead of writing the whole score matrix out and reading it back. That is an implementation detail, not an approximation.

Changing which positions may attend to which, so that less work is done in the first place, is a different idea: sparse attention.

Often misread as:

  • “FlashAttention is an approximation of attention.”

Sparse attention changes the rule of dense attention that every token may attend to every earlier one. In sliding-window attention, each token attends only to the most recent few thousand positions. Other designs combine a local window with a few global tokens, use fixed strided or block patterns, or select positions dynamically.

A token outside the window is not lost. Each layer passes information forward by up to one window, so stacking layers lets information travel much further, indirectly. Several open models alternate local-window layers with full-attention layers to get most of the saving while keeping some direct long-range reach.

Sparse attention is entirely internal to the transformer. It decides which positions already in the context each token may look at, and it is fixed by the architecture. It does not search a document collection, and it does not bring anything new into the context. That is what retrieval does, outside the model.

Often misread as:

  • “Sparse attention fetches the relevant parts of a long document, like search.”

The context window is a limit on sequence length: the largest number of tokens the model can process in one pass. It is not a region of memory inside the model, and not a slice of its embeddings. It is simply the input, measured in tokens.

Everything the model can take into account has to be inside it: the system prompt, the conversation so far, any documents or tool results that were pasted in, and every token the model has generated in this response. When a conversation outgrows it, something has to be dropped or summarised.

The limit comes from several places at once: the lengths the model was trained on, how its positional scheme extends, the quadratic cost of attention, and the memory needed for the KV cache.

Windows have grown from a few thousand tokens to hundreds of thousands and more. That helps, but it does not make retrieval obsolete. Models measurably use information less reliably when it is buried in the middle of a very long input. Every token in the window costs time and money on every request. And most importantly, a window only holds what someone put into it: a document collection of millions of pages still has to be searched, and choosing what to place in the context is what RAG is for.

Often misread as:

  • “The context window is a slice of the model's embeddings.”
  • “Everything inside the context window is used equally well.”
  • “A large enough context window makes retrieval unnecessary.”

A model generates text one token at a time, and each new token attends to every token before it. Without a cache, every step would recompute the keys and values for the entire prefix, again and again.

Because a token’s keys and values never change once computed, the inference engine stores them: for every layer, and for every position processed so far. Each new step computes the query, key and value for the one new token, appends its key and value to the cache, and attends over the stored ones.

The cache grows linearly with the context. Its size is roughly

2×layers×kvheads×headsize×tokens×bytespernumber

which for long contexts can exceed the size of the model weights. Grouped-query attention, quantised caches and paged memory management all exist to keep it in check.

What the cache holds is precise: the attention state for the exact sequence that was actually processed, meaning the prompt plus the tokens that were actually chosen. Candidate tokens that were not sampled were never fed back in, so no keys or values exist for them. The cache is not a store of alternatives.

Nor is it learning. The cache lives for one request, or is kept briefly to reuse a shared prefix such as a long system prompt across requests, which is what “prompt caching” usually refers to. The model’s weights are untouched throughout.

Often misread as:

  • “The KV cache stores alternative completions the model considered.”
  • “The KV cache is a form of long-term memory or learning.”

Choosing the next token

After the last layer, the vector at the final position is multiplied by one more matrix, the unembedding or language-modelling head, which has one column per vocabulary entry. The result is a list of tens of thousands of numbers, one per possible next token. These are the logits.

Logits are raw scores. They can be negative, they do not sum to anything in particular, and only their differences matter: adding the same constant to all of them changes nothing about which token is preferred.

During training the model produces logits at every position at once, and each is compared with the token that really came next. During generation only the last position’s logits are needed, because only the next token is being chosen. Turning these scores into probabilities is the job of softmax.

Often misread as:

  • “Logits are probabilities.”

Softmax exponentiates each logit and divides by the sum, so every token gets a positive probability and all of them add up to one:

pi=ezi∑jezj

It keeps the order of the logits and exaggerates the gaps between them: a logit two points higher than another becomes about seven times as probable. Every token in the vocabulary, however unlikely, keeps a probability above zero.

The distribution describes what tends to come next in text like the context, as learned in training. It is not a calibrated measure of truth. A fluent, confident continuation and a correct one are different things, which is part of why models can state falsehoods with high probability.

The same function appears inside attention, where it turns query-key scores into weights. Same arithmetic, different job.

Often misread as:

  • “A token with high probability is likely to be factually correct.”

The model only ever predicts one next token. Generating a reply is a loop around that:

  1. Run the model on the current sequence and get a distribution for the next token.
  2. Choose one token from it.
  3. Append that token to the sequence and go back to step 1.

The loop stops when the model produces its end-of-sequence token or a length limit is reached. This is why replies stream out token by token: each one really does exist only after the step that chose it.

Greedy decoding always takes the most probable token. Sampling draws a token at random according to the probabilities, usually after reshaping them with temperature, top-k or top-p. Beam search keeps several partial sequences and extends the best few; it is common in translation and rarely used for chat models.

The decoding strategy is part of the inference software, not of the model’s weights. The same model can produce noticeably different text under different settings.

Often misread as:

  • “The model writes a whole answer at once and then displays it word by word.”

These settings act on the final distribution, just before a token is drawn:

  • Temperature divides every logit by a number T before softmax. Below 1 the distribution gets sharper and the likely tokens dominate; as T approaches 0 it approaches greedy decoding. Above 1 it flattens, and unlikely tokens are chosen more often.
  • Top-k keeps only the k most probable tokens and renormalises.
  • Top-p, or nucleus sampling, keeps the smallest set of tokens whose probabilities add up to at least p, so the cut adapts to how spread out the distribution is.
  • Min-p keeps tokens whose probability is at least some fraction of the top token’s.

None of these change what the model computed. They change which of its candidates are eligible and how often each is picked. “More creative” output at a high temperature is the same model drawing more often from the tail of the same distribution.

Even at temperature 0, output is not always perfectly repeatable in practice, because of how floating-point work is batched and parallelised on accelerators.

Often misread as:

  • “Raising the temperature makes the model think more creatively.”

At every step the model produces a probability for every token in its vocabulary, and the decoding loop picks one. What happens to the rest is simple: nothing. They are not fed back into the model, so no hidden states and no KV cache entries are ever computed for them. The sequence continues along the chosen path only.

An API can return the top few probabilities at each step, often called logprobs, but that is a report about one distribution, not a set of alternative continuations.

So “the model considered several answers and rejected them” is misleading for ordinary generation. At each step there was a distribution; after it, there is one token and one sequence. A model can only weigh alternatives in a way that persists if it writes them out as text, which is what reasoning models do when they try an approach, check it and back up, all within one generated sequence.

Speculative decoding looks like an exception but is not. A small draft model proposes several tokens, and the large model verifies them in one pass, keeping the ones it agrees with. Rejected drafts are thrown away, and the result follows the large model’s distribution. It is a speed trick, not a search over answers.

Exploring alternatives deliberately is possible, but it is done by the system around the model: multiple samples and verification.

Often misread as:

  • “During generation the model explores several answers in parallel and keeps the best one.”

If one sample can be wrong, generate several. That simple idea has many forms:

  • Best-of-n: sample n complete answers and pick one with a scoring model or a check.
  • Self-consistency: sample several reasoning paths and take the answer most of them agree on.
  • Search: treat partial solutions as nodes in a tree, extend the promising ones and abandon the rest.
  • Verification: run the unit tests, check the proof, execute the query, and keep what passes.

Each candidate is a separate generation, with its own context and its own KV cache, although a shared prompt prefix can be cached once. The alternatives exist because the inference system asked for them, not because the model kept its discarded candidates. Spending more computation this way, at answer time rather than in training, is often called test-time compute.

None of this changes the model. Candidates that pass a verifier can later be collected and used to train a better model, but that happens afterwards, in a separate process: see synthetic data.

Often misread as:

  • “Running more candidates at inference time trains the model.”

Bringing in outside information

A model knows only what it absorbed in training plus what is in its context. Retrieval-augmented generation adds a step before generation that decides what else should be in that context:

  1. Ahead of time, split a document collection into passages (“chunks”) and index them, usually by computing an embedding for each and often with a keyword index as well.
  2. When a question arrives, embed it the same way and find the passages whose vectors are most similar, optionally re-scoring the best few with a more expensive model.
  3. Insert those passages into the prompt, and let the model generate with them in view.

That gives a model access to private, recent or very large collections without retraining anything, and lets it cite its sources. The model’s weights are not changed at all.

RAG fails in recognisable ways. If retrieval misses the passage that matters, the model answers without it, often fluently. If it retrieves near-misses, they can distract or mislead. Most of the work in a good RAG system is in the retrieval, not the generation.

The Ask view of this article is the retrieval half of this pipeline, without the generating half: it embeds your question in your browser and returns the closest sections, which you then read yourself.

Often misread as:

  • “RAG teaches the model new facts by updating it.”

Both mechanisms compare vectors and favour the most similar, which is why they are easy to confuse. They do different jobs.

Attention runs inside the model, in every layer, for every token. Its queries and keys are the model’s own internal states, computed with weights learned during training. Its scope is exactly the current context window: nothing more, nothing less.

Retrieval usually runs outside the model, before or between generation steps. It is a separate system: a search index, an embedding model and a store of documents. Its scope is whatever collection it indexes, which can be far larger than any context window. Its output is text that gets placed into the context, where attention can then use it.

So sparse attention is not retrieval: it narrows what a token may look at among tokens already present. And a bigger context window does not replace retrieval: it raises the ceiling on how much can be placed in context, but something still has to choose what goes there, and with a large collection it cannot be everything.

A few research architectures build retrieval into the model itself, letting it look up neighbours in a large external store during the forward pass. Even there, the lookup over the store is a separate step from ordinary attention over the context.

Often misread as:

  • “Attention and retrieval are the same thing at different scales.”
  • “A long enough context window makes retrieval unnecessary.”

During ordinary use a model’s weights do not change, and its context is discarded when a request ends. Left to itself, it remembers nothing from one conversation to the next.

Memory features are built around the model instead. The application saves something, such as a summary of a conversation, a note the model was asked to keep, or the files an agent wrote, and later retrieves the relevant parts and places them back into the context. The model “remembers” only by reading.

The storage can be anything: a text file, a database table, a vector index. The hard parts are the same as in RAG: deciding what is worth saving, finding the right piece later, and keeping stored facts accurate as things change.

This is different from the KV cache, which only lasts for one request, and from training, which changes weights in a separate process. Coding agents rely on external memory constantly: task lists, notes, and the repository itself, re-read through tools whenever needed.

Often misread as:

  • “A chat assistant remembers you because the model learned from your earlier conversations.”

When an application describes some tools in the prompt, such as a name, a purpose and the arguments each takes, a model trained for it can respond with a structured call instead of prose: in effect, “call search with these arguments”.

The model executes nothing. The application around it parses the request, decides whether it is allowed, runs it, and appends the result to the context as more tokens. Then generation continues, now with the result in view. Permissions, sandboxing and logging all live in that application, never in the model.

An agent is this loop repeated: generate, call a tool, read the result, generate again, until the task is done. Retrieval is often just one of the tools, alongside running code, reading files or calling an API.

How tools are described to the model and connected to the application varies between vendors. MCP is one attempt to standardise the connection side.

Often misread as:

  • “The model itself executes the code or API calls it requests.”

The Model Context Protocol, introduced by Anthropic in late 2024 and now implemented by many applications, is a standard way to connect AI applications to outside capabilities. A host application, such as a chat app or an editor, runs MCP clients that talk to MCP servers. A server can offer:

  • tools: functions the model may ask to call;
  • resources: data the application can read and place in context;
  • prompts: reusable templates.

Messages are JSON-RPC, carried over standard input and output for local servers or over HTTP for remote ones. The benefit is reuse: an integration written once as an MCP server works in every host that speaks the protocol.

MCP is plumbing. It is not a model, not a memory and not RAG. A server can expose a search tool, and then a retrieval step happens through MCP; but the protocol only carries the request and the result. Ranking, indexing and relevance all belong to whatever the server does, and the model still only sees what ends up in its context. Everything said about tool use still applies: the host decides what runs.

Often misread as:

  • “MCP is a kind of RAG.”
  • “MCP gives the model direct access to systems.”

A repository rarely fits in a context window, and most of it is irrelevant to any given task. So a coding agent does what a developer new to the code would do: it searches, reads, and searches again. The usual instruments:

  • Text search, grep and its descendants: exact and fast, ideal for a known identifier or error message.
  • File listings: the shape of the project, from its directory structure.
  • Symbol indexes and language servers: where something is defined, and everything that refers to it.
  • Embedding search over chunks of code and documentation, for questions phrased in concepts rather than names.
  • Code knowledge graphs: precomputed relations such as calls, imports and ownership, queried as a tool and often exposed over MCP.

The agent chooses which to use, reads the results into context, and lets attention do the reasoning from there. Each tool is a different kind of retrieval, and they complement each other: an exact identifier such as a function name is found reliably by text search and easily blurred by semantic similarity. The Ask view of this article uses the same idea, letting exact keyword matches outrank vague semantic ones.

Often misread as:

  • “A coding agent reads and understands the whole repository at once.”

Learning, and what the words mean

Training is where a model’s weights come from. In pretraining, the model predicts the next token across an enormous body of text; each prediction is compared with the real next token, and gradient descent nudges every weight slightly in the direction that would have made the right token more likely. Post-training then shapes behaviour: supervised fine-tuning on example conversations, learning from human or AI preference judgements, and reinforcement learning on tasks whose answers can be checked.

Inference is everything that happens when the model is used. It runs the network forward only. No error is measured, no gradients are computed, and the weights are read but never written. Inside a conversation, what changes is the context and the KV cache, and both are discarded afterwards.

Models can still adapt to a prompt: show one a few examples of a format and it follows the format. This in-context learning is computation over the context by fixed weights, not training. Remove the examples and the behaviour goes with them.

Whether a provider later uses conversations as training data is a matter of policy, and if it happens, it happens in a separate training run that produces a new model, not inside the conversation.

Often misread as:

  • “A model learns from each conversation as it happens.”
  • “In-context learning updates the model's weights.”

Generated text can be training data. Some common patterns:

  • Generate many solutions to problems with checkable answers, keep the ones that pass, and fine-tune on those.
  • Distillation: train a smaller model on the outputs of a larger one.
  • Have a model critique or rank responses, and use the rankings as preference data.

All of these are an outer loop with a clear sequence: generate, using inference; evaluate, with tests, verifiers or other models; select; then train, which is the only step where weights change, producing a new model. The multi-candidate search of the inference step and the learning of the training step are different stages, often run at different times on different machines.

The loop is only as good as its filter. If verification is weak, errors get trained in and reinforced. Repeatedly training on unfiltered generated text has been shown to narrow a model’s output and lose the rarer parts of the original data, sometimes called model collapse. Where answers can be checked mechanically, as in mathematics and code, the loop works far better than where they cannot.

Often misread as:

  • “A model improves itself every time it generates and checks an answer.”

The hypothesis is old: a system able to improve the system that produces it could make each improvement come faster than the last. In 1965 the statistician I. J. Good called the possible result an “intelligence explosion”.

Pieces of today’s practice point in that direction. Models generate and filter their own training data, write code for machine learning experiments, and help evaluate other models. But every one of those loops still runs through training jobs, compute budgets and evaluations that people design and launch, and no model changes its own weights while it runs.

Whether such loops could become autonomous and compounding, or would run into limits of data, compute, or the ability to verify that a change really is an improvement, is an open research question. The honest summary is that the concept is well defined and seriously studied, and the mechanism does not exist in deployed systems today.

Often misread as:

  • “Current models rewrite their own weights to become smarter while they run.”

“Artificial general intelligence” is used in at least four quite different senses:

  • Human-level breadth: performing at or above human level across most cognitive tasks, not only one domain.
  • Learning ability: acquiring new skills as efficiently as a person does, from little data.
  • Economic: doing most economically valuable work. OpenAI’s charter uses a version of this, describing systems that “outperform humans at most economically valuable work”.
  • Graded: frameworks that drop the single threshold and rate systems by levels of performance and generality instead.

The disagreement is not just pedantry. “General” and “intelligence” are both contested words, the definitions point to different measurements (benchmark scores, job tasks, autonomy, learning speed), and organisations have commercial and contractual reasons to prefer one or another.

So the useful question about any AGI claim is not yes or no, but: which definition, measured how? A system can satisfy one definition comfortably while clearly failing another.

Often misread as:

  • “AGI is a well-defined threshold that a model either passes or does not.”

Looking for one idea in particular? Ask this article finds the closest sections using the same kind of retrieval described in Retrieval-augmented generation, running on your device.