From the «AI agents» series

70 terms

AI agents glossary

Agentic development terms in plain words: definition, example, where it shows up and the primary source. Each term is also a «Term of the day» card in our Telegram channel.

A

Agent fleet

fleet · multi-agent

Several agents working in parallel, each in its own lane: subagents, parallel sessions, background and cloud agents, teams run by an orchestrator.

Agent memory

auto memory · memory

How knowledge carries across sessions: instruction files you write and notes the agent keeps itself from your corrections.

Agent SDK

Claude Agent SDK · OpenAI Agents SDK

A library for building your own agent: a ready loop, tools, context management and permissions, so you can embed an agent in your product.

Agent teams

agent teams

Several independent sessions run by a lead, with a shared task list and messages between members. Each has its own context and can be addressed directly.

Agent-computer interface

ACI

Everything through which an agent works with a system: the tools, their names, descriptions, parameters and responses. It deserves as much design care as a human interface.

Agentic loop

agentic loop · agent loop

The cycle an agent works through: gather context, take action, verify results, and repeat until the task is done.

AGENTS.md

AGENTS.md

An open file format of instructions for coding agents, supported by Codex, Cursor, Copilot, Gemini CLI and many others; the nearest AGENTS.md takes precedence.

AI agent

agent · агент

A system where a language model decides which steps to take and which tools to use, and tracks how the task is going on its own.

Attention

attention · self-attention

The mechanism by which a model decides, for each token, which other tokens in the context to rely on and how much. The core of the transformer.

B

Background agent

background agent · agent view

An agent that works while you do something else and only calls you when it needs an answer, without blocking the main conversation.

Benchmark

benchmark · SWE-bench

A public, standardized task set used to compare models and agents. Unlike your own evals, it reflects typical tasks, not yours.

C

Checkpoint

checkpoint · /rewind

A restore point the agent creates before making changes: you can roll back the code, the conversation, or both to it.

CLAUDE.md

CLAUDE.local.md

A file of persistent instructions for Claude Code: build and test commands, architecture, project rules. Loaded at the start of every session.

Cloud agent

cloud agent · Cloud Agents

An agent that works in its own virtual machine with a clone of the repo and keeps going without your computer. The result usually arrives as a pull request.

Compaction

compaction · /compact · сжатие контекста

Automatic summarization of the conversation when the context window nears its limit: old tool output is cleared, the rest is condensed.

Computer use

computer use

A tool through which the model sees the screen via screenshots and controls the mouse and keyboard, working with interfaces that have no API.

Context engineering

context engineering

Curating what the model sees when it decides: picking and maintaining the smallest set of high-value tokens — instructions, tools, data, history.

Context rot

context rot

The decline in a model's performance as the context grows: it has a limited attention budget, and every new token spends some of it.

Context window

context window · контекст

The model's working memory: all the text it takes into account when answering — instructions, history, files, command output and the reply itself. Measured in tokens.

Cursor rules

.cursor/rules · .mdc

Instructions for Cursor's agent: .mdc files in .cursor/rules with description, globs and alwaysApply fields, applied always, intelligently, by file pattern or manually.

E

Embeddings

embeddings · векторные представления

Numeric representations of text that let you measure semantic similarity, the basis of semantic search, recommendations and RAG.

Eval

evals · оценка · evaluation

A test for an AI system: give the agent a task and grade the result with predefined logic. A set of them shows whether a change made things better.

Evaluator-optimizer

evaluator-optimizer

A pattern: one model generates a response, another evaluates it and sends feedback, and the loop repeats until the evaluator is satisfied.

Excessive agency

excessive agency · LLM06

A vulnerability where an agent has too many functions, permissions or freedom to act without human checks, so a model error or injection causes real damage.

Extended thinking

extended thinking · adaptive thinking · effort

The model's step-by-step reasoning before it answers. In recent models the model decides how deep to go, steered by an effort level.

F

Fine-tuning

fine-tuning · файнтюнинг

Further training of an existing model on your own examples so it handles a narrow task or format better. Unlike RAG, it changes the model itself rather than its context.

G

Git worktree

worktree · --worktree

An extra working copy of the repo in its own folder on its own branch, letting several agents work in parallel without overwriting each other's files.

Guardrails

guardrails

Checks on an agent's input and output, often run on a fast, cheap model: is the agent being asked for something off-limits, does the reply contain what it shouldn't.

H

Hallucination

hallucination · конфабуляция

A confident, plausible but wrong model answer: a function that doesn't exist, a made-up number or citation. The model doesn't “know” it's wrong.

Handoff

handoff

A mechanism where one agent hands a task to another, better-suited one, which takes over the conversation.

Harness

harness · agent harness · агентная обвязка

The scaffolding that turns a model into an agent: the loop, tools, context management, permissions and execution environment. The model thinks; the harness gives it hands and rules.

Hook

hook · PreToolUse · PostToolUse

A handler that fires at a fixed point in the agent's loop: before a tool call, after a file edit, at the end of a turn. Unlike instructions, it always runs.

Human in the loop

human-in-the-loop · HITL

A setup where the agent stops and waits for a human decision before consequential actions: writes, deletes, payments, sending.

I

Inference

inference · вывод модели

Using an already trained model: a prompt in, a response out. Inference is billed by tokens and drives an agent's latency and cost.

L

Least privilege

least privilege · принцип минимальных привилегий

A security principle: the agent and its tools get exactly the access the task needs, and nothing more.

LLM

large language model · большая языковая модель

A large language model: a neural network trained on vast amounts of text that can write text and code, answer questions, summarize and reason.

LLM as a judge

LLM-as-judge · model-based grader

Grading an agent's output with another model against a rubric. More flexible than code, cheaper than a human, but it needs checking itself.

LoRA

Low-Rank Adaptation

A cheap way to fine-tune: the model's weights stay frozen and small low-rank matrices are trained alongside them, cutting trainable parameters for GPT-3 by 10,000x, per the authors.

M

MCP

Model Context Protocol

Model Context Protocol: an open protocol that standardizes how LLM apps connect to external data and tools, one format for a tracker, a database, a browser and any service.

MCP host and client

MCP host · MCP client

The host is the LLM app that connects to MCP servers (Claude Code, Codex, Cursor). A client is the connector inside the host that talks to one server.

MCP server

MCP server

A program that gives an agent tools, resources (data) and prompts over MCP. It can be local (stdio) or remote (Streamable HTTP).

Multimodality

multimodal · multimodal model

A model's ability to handle several kinds of data — text, images, audio, video — as input, output or both.

N

Non-interactive mode

headless · claude -p · codex exec

A mode where the agent runs a single task and exits without a dialog — the basis for scripts, CI and automation.

O

Orchestrator

orchestrator-workers · оркестратор и исполнители

A lead model that splits a task into parts, hands them to workers and combines the results. The orchestrator-workers pattern underlies multi-agent systems.

P

Parallelization

parallelization · sectioning · voting

A pattern: independent parts run at the same time, or one task is solved several times and the results are compared by voting.

pass@k and pass^k

pass@k · pass^k

Two reliability metrics. pass@k is the chance of at least one success in k tries; pass^k is the chance that all k tries succeed.

Permission mode

permission mode · auto mode · acceptEdits

The session's baseline autonomy: what the agent does without asking. In Claude Code: plan, default (Manual), acceptEdits, auto, dontAsk and bypassPermissions.

Permission rules

permissions · allow · deny

Rules that allow, ask about or deny agent actions by tool name and argument pattern, evaluated deny → ask → allow.

Plan mode

plan mode

A mode where the agent studies the code and proposes a plan without changing files. You approve the plan before implementation starts.

Plugin

plugin · marketplace

A bundle of skills, hooks, subagents and MCP servers installed as one unit and distributed to a team via a marketplace.

Pretraining

pre-training · pretraining

The first and most expensive training stage: the model learns to predict the next token over a huge text corpus and gains general knowledge of language and the world.

Prompt chaining

prompt chaining

A workflow pattern: the task is split into sequential steps, each model call processes the previous output, with checks between steps.

Prompt engineering

prompt engineering

The craft of writing requests and instructions so that the model's output is predictable: the task, context, output format, examples, criteria.

Prompt injection

prompt injection · LLM01

An attack where someone else's instructions hidden in a prompt, file, web page or tool result steer the model away from your task. First in the OWASP Top 10 for LLMs.

Q

Quantization

quantization · квантование

Storing model weights at lower precision, say 8 or 4 bits instead of 16. The model takes less memory and runs faster at a small cost in quality.

R

RAG

retrieval-augmented generation · генерация с поиском

A technique where relevant passages are retrieved from a knowledge base and put into the model's context before it answers, grounding the reply in facts.

RLHF

reinforcement learning from human feedback · обучение с подкреплением на отзывах людей

Reinforcement learning from human feedback: people compare or rate the model's answers, and the model is trained on those ratings to answer more helpfully and safely.

Routing

routing

A workflow pattern: the input is classified and sent to a specialized branch with its own prompt, tools or model.

S

Sandbox

sandbox · sandboxing

An OS-enforced boundary on which files and network hosts the agent's commands can reach. It works independently of permission rules.

Skill

skill · SKILL.md · Agent Skills

A folder with a SKILL.md file: a description of when to use it, plus instructions. It loads only when needed — by the agent or via /name.

Slash command

slash command · /command

A command invoked with / in the agent's prompt. Built-in ones control the session; your own package repeatable routines.

Subagent

subagent · sub-agent

A helper with its own context window, system prompt, tools and permissions. It takes a subtask and returns only the result to the main conversation.

System prompt

system prompt

Instructions the app sends the model ahead of the conversation: the role, rules, how to use tools and format replies.

T

Token

token

The smallest unit of text for a model: a word, part of a word or a character. Context windows, limits and pricing are all counted in tokens.

Tool use

tool use · function calling

A model's ability to call functions: it decides when a tool is needed and returns a structured call that the app or the provider executes.

Transcript

transcript · trace

The complete record of an agent run: messages, tool calls, reasoning, intermediate results. It shows why the agent did what it did.

Transformer

transformer

The neural network architecture behind modern LLMs. It processes a sequence of tokens through layers of attention, without recurrent or convolutional networks.

V

Vector database

vector database · vector store · pgvector

A store for embeddings with fast nearest-neighbour search. It finds passages close in meaning to a query — the backbone of RAG.

Verification loop

verification loop

A check the agent runs itself — tests, a build, a screenshot — iterating until it passes instead of stopping after one attempt.

W

Workflow

workflow · рабочий процесс

A system where the model and tools follow steps defined in code in advance. Unlike an agent, the model doesn't choose what to do next.