HACKS VITAE

TOOLS & TECH · SOURCES SHOWN

How Large Language Models Work, in Plain Words

PRICES, VERSIONS AND FACTS AS OF SEPTEMBER 2026

ChatGPT, Claude, Gemini and similar AI chatbots run on large language models. Here is what those models do — tokens, next-word prediction, training, the extra steps that turn them into assistants, reasoning models and knowledge cutoffs — explained from the labs' own papers.

9 SECTIONS · HOVER A POINT TO JUMP
Published
September 17, 2026
Facts as of
September 2026
Read
7 min
Sections
9

BACKGROUND · VELÁZQUEZ, JUAN DE PAREJA, 1650 · THE MET, OPEN ACCESS

THE SHORT VERSION

  1. A large language model is first trained to predict the next piece of text. OpenAI describes GPT-4 as "pre-trained to predict the next token in a document".
  2. It reads and writes in tokens, not words. In English a token is roughly three-quarters of a word, but the ratio depends on the model.
  3. A model trained only to predict text is not yet a helpful assistant. Extra training, for example using human or AI feedback, turns it into one.
  4. Two practical consequences: the same question can get different answers, and a model's built-in knowledge stops at a cutoff date unless it can search.

THE ARTICLE · 7 MIN

AI chatbots such as ChatGPT, Claude and Gemini are built on large language models. You do not need maths to understand the basic idea, and knowing it explains a lot: why answers vary, why a model can sound confident and still be wrong, and why it may not know about recent events. This page explains the main steps in plain words, using the labs’ and researchers’ own descriptions. Details change quickly; this is a snapshot as of September 2026.

1. It predicts the next piece of text

The models behind these chatbots start from one basic task: given some text, predict what comes next. OpenAI’s technical report describes its model this way: “GPT-4 is a Transformer-based model pre-trained to predict the next token in a document.”

Anthropic’s glossary gives the same idea in plainer words: models “are pretrained to predict the next word, given the previous context of text in the document”. To write a long answer, models like these repeat that step: predict a piece, add it to the text, predict the next piece.

2. It works in tokens, not words

The “pieces” are called tokens. OpenAI’s help centre explains: “A token can represent a character, part of a word, a whole word, or punctuation.” It is direct about the difference: “A token count is not the same as a word count.” For English, it gives a rough guide: “1 token is approximately three-quarters of a word. 100 tokens are approximately 75 words.”

That ratio is not fixed. Anthropic notes that its Claude 4.7 and later models use a newer tokenizer and that “This tokenizer produces approximately 30% more tokens for the same text.” OpenAI adds that “The same text can produce different token counts depending on the model, its encoding, and the language.” Tokens matter to you because model limits and costs are counted in them.

3. The Transformer: paying attention to the whole sentence

The GPT-4 report above calls the model “Transformer-based”. The Transformer design was introduced in the 2017 paper “Attention Is All You Need”, which proposed an architecture “based solely on attention mechanisms, dispensing with recurrence and convolutions entirely”. It was tested on translation, including an “English-to-German translation task”.

Google’s blog post about it explained the key idea, self-attention, which “directly models relationships between all words in a sentence, regardless of their respective position.” Its example was the sentence “I arrived at the bank after crossing the river”: to know that “bank” means a riverbank, the model has to connect it with “river”, several words away.

4. Training: a very large amount of text

A model learns by adjusting a huge set of internal numbers, called parameters, until its predictions improve. Model sizes and, above all, the amount of training text have grown fast:

ModelParametersTraining text
GPT-3 (2020)175 billion — “an autoregressive language model with 175 billion parameters”“All models were trained for a total of 300 billion tokens.”
Chinchilla (2022)70 billion“We verify this by training a more compute-optimal 70B model, called Chinchilla, on 1.4 trillion tokens.”
Llama 3 (2024)405 billion — “405B trainable parameters on 15.6T text tokens”15.6 trillion tokens

Researchers found that this kind of growth pays off in a predictable way. A 2020 study reported that “The loss scales as a power-law with model size, dataset size, and the amount of compute used for training”. “Loss” here is the paper’s measure of prediction error (cross-entropy loss), not a direct score of what a model can do. A 2022 study then concluded that “current large language models are significantly undertrained” and that, “for compute-optimal training”, “for every doubling of model size the number of training tokens should also be doubled”.

Some of the newest model descriptions give no total: OpenAI says GPT-6 Astra “was trained on diverse datasets”, and Anthropic says Claude Opus 5 “was pretrained on large, diverse datasets”. Some open-weight developers still publish a figure: DeepSeek’s report says “we train DeepSeek-V4-Flash on 32T tokens and DeepSeek-V4-Pro on 33T tokens”.

5. From text predictor to assistant

A model trained only to predict text is not yet a chatbot. Anthropic’s glossary says: “These pretrained models are not inherently good at answering questions or following instructions”. A further round of training shapes it into an assistant.

One method, used for both GPT-4 and Claude, is reinforcement learning from human feedback (RLHF). In OpenAI’s InstructGPT work, people first wrote “demonstrations of the desired model behavior”. The paper continues: “We then collect a dataset of rankings of model outputs, which we use to further fine-tune this supervised model using reinforcement learning from human feedback.” The result was striking but limited in scope: “In human evaluations on our prompt distribution, outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters.” The same paper noted that “InstructGPT still makes simple mistakes”.

Anthropic’s Constitutional AI trained a harmless assistant “without any human labels identifying harmful outputs”: “The only human oversight is provided through a list of rules or principles”.

6. Reasoning models: thinking before answering

Some models are now trained to work through a problem step by step before giving an answer. OpenAI’s system card for its o1 models says: “The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought.” In a post dated September 12, 2024, OpenAI also reported that performance improved “with more time spent thinking (test-time compute)”.

DeepSeek’s R1 paper argued that “the reasoning abilities of LLMs can be incentivized through pure reinforcement learning (RL), obviating the need for human-labeled reasoning trajectories”. Thinking also costs tokens. OpenAI’s help centre says reasoning tokens “are not visible as answer text, but they count toward output usage and are billed as output tokens.”

7. Why the same question can get different answers

When a model picks its next token, there can be an element of chance. Anthropic’s glossary describes temperature as “a parameter that controls the randomness of a model’s predictions during text generation”, and explains: “Lower temperatures result in more conservative and deterministic outputs that stick to the most probable phrasing and answers.” Even so: “Even with temperature set to 0, the results will not be fully deterministic and identical inputs may produce different outputs across API calls.”

Some newer models do not let developers change it at all. Anthropic’s API documentation says: “Models released after Claude Opus 4.6 do not support setting temperature.”

8. Why a model may not know recent events

A model’s built-in knowledge comes from its training text, which stops at some point. Labs sometimes give two different dates. Anthropic says its Claude Opus 4 and Claude Sonnet 4 models have a training data cutoff of March 2025 but a reliable knowledge cutoff of January 2025, “which means the models’ knowledge base is most extensive and reliable on information and events up to January 2025”.

To get past that limit, models can be connected to tools such as web search. Meta’s Llama 3 paper says the model was trained to use a search engine “to answer questions about recent events that go beyond its knowledge cutoff”.

What this explains

  • Fluent is not the same as checked. The model is trained to produce likely text; predicting the next token does not, by itself, check that text against a source. The GPT-4 report says the model “still is not fully reliable (it “hallucinates” facts and makes reasoning errors)”.
  • Answers vary. Randomness in generation means the same question can produce different wording, and sometimes a different answer.
  • Recent events are a weak spot unless the model is searching, and then the answer depends on what it found.

Our reading: a large language model is best thought of as an extremely capable text predictor that has been further trained to be helpful. That makes it very useful for drafting, explaining and summarising — and it is exactly why any fact that matters should be checked against a source.

Sources

Checked September 2026.

Related: What is a frontier AI model? · AI benchmarks explained · Open-weight is not open-source

  • artificial intelligence
  • ai models
  • large language models
  • explainer

SHARE & CITE

Hacks Vitae. "How Large Language Models Work, in Plain Words." September 17, 2026. https://www.hacksvitae.com/life-hack/how-large-language-models-work-in-plain-words

That's what we found. The rest is your call.

118 articles, each with its sources listed. Spotted something off? hacksvitae@gmail.com

Open the library