
Agents, RAG, evals, MCP, orchestration… This vocabulary lands in your product meetings without its definition. Running an AI project doesn't require you to code. This vocabulary does: it's what lets you make the call, and ask the question that separates a real project from a prototype.
21 terms, in three parts: the building blocks, the tooling, quality.
The building blocks
LLM (Large Language Model)
The engine that generates text from whatever you give it. The "brain" behind Claude, ChatGPT and the rest.
Example Two teams on the same model get opposite results: the one that feeds it their specs, tickets and business rules, and the one that asks the question with nothing attached.
What it changes An LLM is a component you have to equip. The value sits in what you put around it: the context, the tools, the guardrails.
Multimodal
A model that handles text, images and voice.
Example In support, the user sends a screenshot of the error instead of trying to describe it. Triage happens on the image.
What it changes New use cases open up: reading a screenshot, a scanned document, a recording.
Prompt
The instruction you give the model for one specific task.
Example "Summarise this ticket" and "Summarise this ticket for a developer who hasn't followed the thread, in three points, one being user impact" don't produce the same output.
What it changes Framing a problem clearly becomes a core product skill. Getting the framing right is already half the result.
System prompt
The standing instruction placed above every conversation: the model's role, its rules, what it must refuse.
Example "Don't propose a technical solution before stating the problem" is written once in the system prompt, and applies to every session across the team.
What it changes It's what makes results consistent from one person to the next. Without it, everyone reinvents their own instructions and gets something different.
Context (context engineering)
Everything you hand the model for a task: documents, history, rules.
Example An agent that doesn't know your internal naming writes specs that are correct on substance and unusable as they stand.
What it changes What you give the model weighs more than which model you pick.
Tokens & context window
A token is the unit that measures length and cost; the context window is what the model can hold in mind at once.
Example An agent that re-reads your entire documentation on every question can cost ten to twenty times more than one that fetches the three relevant paragraphs.
What it changes AI cost follows usage, token by token. Agentic use can blow it up — budget it as a variable line.
Reasoning (chain-of-thought)
Models that "think" step by step before answering.
Example Weighing three roadmap options with their dependencies warrants it. Rewording a ticket title doesn't.
What it changes More reliability on complex tasks, at the price of more tokens and more time. Worth reserving for the cases that justify it.
Non-determinism
Given the same question, the model can return a different answer.
Example The same ticket submitted twice, two minutes apart, comes back as P1, then P2. Both answers hold up.
What it changes You can't promise an identical output on every run. What you steer is an acceptable range.
Hallucination
When the model invents a wrong answer, and states it with confidence.
Example An API version number that never existed, quoted among four accurate references, in the same tone.
What it changes A wrong answer doesn't announce itself. Anywhere an error is expensive, decide who checks, and what they check.
The tooling
Knowledge base (RAG, Graph-RAG…)
Giving the system access to your own content so it answers from it. Several techniques coexist: RAG fetches the relevant passages, Graph-RAG also uses the links between entities, other approaches index whole documents.
Example The support agent answers from your product documentation and your resolved tickets, on the version actually running at your company.
What it changes This is what gives you an AI that knows YOUR business. Picking the technique is an engineering call; what's yours is the quality and freshness of what goes in.
Fine-tuning
Retraining a model on your data to specialise its behaviour.
Example An image model retrained on your art direction produces visuals in your brand style. For business text, a knowledge base gets you there for a fraction of the cost.
What it changes Powerful, expensive, and it calls for AI skills few product teams have in house. The useful reflex is to start with context, and only consider fine-tuning if the model's behaviour is still the problem.
MCP (Model Context Protocol)
The standard that describes how a tool plugs into the AI: Gmail, Notion, your database.
Example Switching ticketing tools doesn't mean rewriting the agent. You swap a connector, everything else keeps working.
What it changes It's the standard API for agents. Writing in-house connectors in 2026 burns budget.
Tool use
The model's ability to trigger an action itself, and read the result. Where MCP describes how the tool plugs in, tool use is the moment the model reaches for it.
Example You ask about the state of a release. The model reads the tickets, cross-checks the last deployment, and answers. Nobody opened a tab.
What it changes This is the shift from an assistant that drafts to an agent that executes. Any task that ends inside a tool becomes delegable.
Persistent memory
The system keeps what it learned from your exchanges across sessions: conventions, preferences, past decisions.
Example After three exchanges where you correct the same thing, the instruction sticks. The fourth time, you don't give it.
What it changes The AI stops being reset at every conversation. You stop re-explaining context, and a correction made once holds.
Skill
A packaged, reusable capability you hand the AI: know-how written down.
Example Your method for writing a PRD, captured once, becomes usable by the whole product team with the same result.
What it changes Skills turn a personal habit into a repeatable team process. It's the bridge between "I tinker" and "we industrialise".
Agent
An AI that chains several steps on its own to reach a goal, rather than answering in one shot.
Example Preparing the weekly review: the agent reads the closed tickets, spots the gaps against the roadmap, and produces the summary note. A dozen steps, no prompting in between.
What it changes Long or tedious tasks become automatable end to end: a full sequence, rather than one answer at a time.
Orchestration
Coordinating several agents or steps: deciding who does what, and in which order.
Example A lead agent hands out the work: one drafts the spec, a second challenges it on edge cases, a third checks that the figures quoted exist. It arbitrates and returns the consolidated result.
What it changes The new core skill is orchestrating and reviewing.
Quality
Evals
How you measure whether the AI does its job, against objective criteria.
Example Thirty real tickets, with the right answer annotated by hand, replayed on every prompt change. The score compares one version to the next.
What it changes These are your acceptance tests. Without a test set, you can neither compare two versions nor show that a change improves anything. The question to raise in a steering meeting: who writes the evals, and who reviews them?
Guardrails
The limits you set to keep the AI from going off the rails: security, scope, tone.
Example The agent may read your database and propose a change. It may not write.
What it changes Setting guardrails forces you to write down what the system isn't allowed to do. Few teams formalise it, even though it's often the easiest part to settle in a meeting.
Human-in-the-loop
Keeping a human who validates at the critical steps.
Example The agent prepares the reply to the customer and leaves it as a draft. A human sends it.
What it changes The working rhythm changes. The agent produces continuously, the human validates in batches: you need to know the agents' cycles to place checkpoints without creating a queue.
Vibe coding
Coding "on feel" with AI, with no spec and no frame.
Example The prototype ships in two days and wins the room. Four months later, nobody knows why it works, or what breaks when you touch it.
What it changes You gain velocity on design and prototyping: an idea becomes testable in hours. In exchange, almost nothing is documented — a trade-off to own, depending on how long what you're building has to live.
What this is really about
AI reshuffles the deck inside product teams. A PM can now step outside their own patch: ship a prototype, a mockup, an analysis that used to need someone else's expertise. That still takes understanding AI and its limits, to keep control of what you ship, move faster on better assets, and pick up new skills along the way.
It also takes change management in step with the teams next door. Design, engineering and data are watching their own scope widen: the boundaries are being renegotiated from every side at once.
This is exactly what I work on with product teams. If it's a live topic at your company, let's talk.


