THE ARTICLE · 7 MIN
AI news is full of jargon. This glossary explains 22 common terms in plain words. Most entries quote a short definition from a named source — usually Google’s machine learning glossary, an AI company’s own documentation, a research paper or the Open Source Initiative — followed by what it means in practice. Usage changes quickly; this reflects September 2026.
For the bigger picture, see how large language models work.
The basics
1. Large language model (LLM)
Google’s glossary: “At a minimum, a language model having a very high number of parameters.” It adds that, more informally, the term covers “any Transformer-based language model, such as Gemini or GPT”.
In practice: the AI system behind a chatbot. The chat app is the product; the model is the engine inside it. Anthropic, for example, describes its assistant Claude as one that “is based on a large language model that has been fine-tuned and trained using RLHF”.
2. Token
Google’s glossary: “In a language model, the atomic unit that the model is training on and making predictions on.”
In practice: models read and write in tokens, and usage limits and prices are counted in them. How much text fits in a token varies. OpenAI’s help centre says “1 token is approximately 4 characters” for English text, while Anthropic’s glossary says that for Claude “a token approximately represents 3.5 English characters, though the exact number can vary depending on the language used.” Anthropic has also said that Claude Opus 4.7’s updated tokenizer can turn the same text into “roughly 1.0–1.35×” as many tokens as Claude Opus 4.6, depending on the content type, so any such figure is approximate.
3. Context window
Google’s glossary: “The number of tokens a model can process in a given prompt.”
In practice: how much text — your question, pasted documents, the conversation so far and the model’s own reply — the model can handle at once.
4. Parameters
Google’s glossary: “The weights and biases that a model learns during training.”
In practice: the internal numbers that training adjusts. A bigger parameter count means a bigger model, but not automatically a better one.
5. Mixture of experts (MoE)
Google’s glossary: “A scheme to increase neural network efficiency by using only a subset of its parameters (known as an expert) to process a given input token or example.”
In practice: this is why some models quote two sizes. DeepSeek’s V3 report, for example, describes a model “with 671B total parameters with 37B activated for each token”.
6. Training compute (FLOP)
Epoch AI: “Compute usage is commonly measured as the number of floating point operations (FLOP) required to train the final version of the system”.
In practice: a count of the calculations used to train a model. Some laws use it as a threshold — New York’s definition of a “frontier model” uses 10^26 operations; see what is a frontier AI model?
Training and using a model
7. Pretraining
Google’s glossary: “The initial training of a model on a large dataset.”
In practice: the first stage of training, on a very large amount of text. Google adds that “Some pre-trained models are clumsy giants and must typically be refined through additional training.”
8. Fine-tuning
Google’s glossary: “A second, task-specific training pass performed on a pre-trained model to refine its parameters for a specific use case.”
In practice: extra training that shapes a general model for a purpose. One common kind is instruction tuning, which Google describes as “A form of fine-tuning that improves a generative AI model’s ability to follow instructions.”
9. RLHF (reinforcement learning from human feedback)
Google’s glossary: “Using feedback from human raters to improve the quality of a model’s responses.”
In practice: people compare or rank model answers, and the model is trained towards the preferred ones. OpenAI’s InstructGPT paper, which combined supervised fine-tuning with RLHF, reported that “outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters” in its human evaluations.
10. Inference
Google’s glossary: “In large language models, inference is the process of using a trained model to generate a response to an input prompt.”
In practice: using the model, as opposed to training it. Google notes the word “has a somewhat different meaning in statistics”.
11. Temperature
Google’s glossary: “A hyperparameter that controls the degree of randomness of a model’s output.”
In practice: one reason the same question can get different answers. Some newer models no longer let developers set it.
12. Reasoning model
OpenAI: “Reasoning models use internal reasoning tokens before producing a response.” Google uses the word “thinking”: “When you use a thinking model, Gemini reasons internally before responding.”
In practice: models that work through a problem before answering. Companies use different names for the same idea, and the extra thinking uses tokens.
What models can do
13. Multimodal
Google’s glossary: “A model whose inputs, outputs, or both include more than one modality.”
In practice: a model that handles more than text, such as images, audio or video.
14. Tool use (function calling)
Anthropic: “Tool use (also called function calling) lets Claude call functions that you define or that Anthropic provides.” The model “returns a structured call that your application executes”.
In practice: the model does not run the tool itself. It asks for a tool — a search, a calculator, a database lookup — and software carries out the request.
15. Agent
There is no single definition. Google’s glossary: “Software that can reason about user inputs in order to plan and execute actions on behalf of the user.” Anthropic draws a line between workflows, where “LLMs and tools are orchestrated through predefined code paths”, and agents, “systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks”.
In practice: a system that takes several steps on its own, often using tools, rather than giving one answer.
16. Hallucination
Google’s glossary: “The production of plausible-seeming but factually incorrect output by a generative AI model that purports to be making an assertion about the real world.”
In practice: a plausible-sounding answer that is wrong. The word is debated — see why AI makes things up.
17. Knowledge cutoff
In practice: the point after which a model’s training data stops, so it may not know about later events unless it can search. The term can refer to more than one date. Anthropic, for example, gave two dates for its now-retired Claude Opus 4 and Claude Sonnet 4: a “training data cutoff date of March 2025” and “a reliable knowledge cutoff date of January 2025”.
Openness and documentation
18. Open weights
The Open Source Initiative: “Open Weights refer to the final weights and biases of a trained neural network.”
In practice: the trained weights are published, so you can download and run the model yourself, under whatever licence comes with them. The same page notes that “Open Weights differ significantly from Open Source AI”.
19. Open-source AI
The Open Source Initiative’s definition: “An Open Source AI is an AI system made available under terms and in a way that grant the freedoms” to use, study, modify and share it. It requires detailed information about the training data, “The complete source code used to train and run the system”, and the model parameters.
In practice: a much higher bar than downloadable weights. See open-weight is not open-source.
20. Model card
From the 2018 paper that proposed them: “Model cards are short documents accompanying trained machine learning models that provide benchmarked evaluation in a variety of conditions”.
In practice: a document describing a model, how it was evaluated and the context it is intended for.
21. System card
There is no standard definition. Anthropic describes its own: “System cards document the capabilities, safety evaluations, and responsible deployment decisions for Claude models.”
In practice: AI companies use the term for a longer report published with a model, often focused on safety testing. The name has also meant other things: Meta used it in February 2022 for “a new resource for understanding how AI systems work”.
22. Benchmark
Google’s glossary defines LLM evaluations as “A set of metrics and benchmarks for assessing the performance of large language models”.
In practice: a standard set of questions or tasks, scored the same way for every model, used to compare models on the same test. Scores depend on how the test is run — see AI benchmarks explained.
Sources
- Google for Developers, Machine Learning Glossary. Definitions quoted from it are © Google, licensed under CC BY 4.0.
- Anthropic, Claude Platform Docs: Glossary; Tool use overview; Anthropic, “Building effective agents” (December 2024); Transparency Hub; System cards.
- OpenAI, API docs: Reasoning models; Help Center, “Understanding and counting tokens”.
- Anthropic, “Introducing Claude Opus 4.7”; Meta, “System Cards, a new resource for understanding how AI systems work” (23 February 2022).
- New York Senate Bill S8828 (2026), frontier model definition.
- Google, Gemini API docs: Thinking.
- DeepSeek-AI, “DeepSeek-V3 Technical Report”, arXiv:2412.19437.
- Epoch AI, “Estimating Training Compute of Deep Learning Models” (2022).
- Ouyang et al., “Training language models to follow instructions with human feedback”, arXiv:2203.02155.
- Open Source Initiative, “Open Weights”; “The Open Source AI Definition – 1.0”.
- Mitchell et al., “Model Cards for Model Reporting”, arXiv:1810.03993.
Checked September 2026.
Related: How large language models work · AI myths checked · Who makes today’s leading AI models?
- artificial intelligence
- ai models
- glossary
- explainer
