Jonas Hansen

Ask this article

Ask a question in your own words, and get back the sections of How modern language models work that answer it best. This is retrieval, not a chatbot: it finds sections, it does not write answers.

How this works

This page is a small retrieval system of the kind the article describes in Retrieval-augmented generation. Every step happens in your browser; your question is never sent anywhere.

  1. Tokenize. Your question is split into the WordPiece tokens of the model's vocabulary (tokenization).
  2. Embed. Each token selects one row of a static embedding table, and the rows are averaged into a single vector of 128 numbers (embeddings).
  3. Compare. That vector is compared with vectors computed for every passage of the article when the site was built, by cosine similarity.
  4. Match words. Exact words and phrases from your question are matched against each section's title, concepts and text, so a precise term like KV cache is never outranked by something only vaguely similar.
  5. Rank. The two signals are combined, and the best sections are listed above, each linking to its place in the article.

Model: potion-base-4M by Minish Lab (MIT licence), a static embedding table of 128 dimensions, quantised to 8 bits. It is served from this site and downloaded once, when you first use the form.