Post 6 of 8
What you need to have an AI agent of your own
The seven parts of a working agent: model, harness, tools, context, checks, permissions and budget. What you can take off the shelf, what you'll have to build and what it costs in tokens.
Alexander Mihalkevich · Fact-checked October 5, 2026 · 5 min read
“I want my own agent” usually means one of two things: an agent that helps you work with code and documents, or an agent inside your product that works with customers and data. Their components are the same; the difference is what you take off the shelf. Below are seven components and a minimal starter set.
1. The model
The model reasons and chooses the next action. Access is by subscription or API. Claude Code requires a Pro, Max, Team or Enterprise subscription or a Console account; the free claude.ai plan does not include access. You can also work through cloud providers (Amazon Bedrock, Google Cloud, Microsoft Foundry).
According to the models overview, Claude Fable 5.1, Opus 5.5 and Sonnet 5.5 have a 1M-token context window, and Haiku 4.5 has 200K. A large window is no reason to load everything into it; more on that below.
2. The harness
A harness is the scaffolding around the model: the “gather context → act → verify” loop, tools, context management and permissions. The Claude Code documentation puts it plainly: Claude Code is the harness, Claude is the model inside it.
- An off-the-shelf harness: Claude Code, Codex, Cursor. Suited to working with code, documents and data on your machine or in the cloud.
- Your own harness: built on the Claude Agent SDK (Python and TypeScript; the same engine as Claude Code) or the OpenAI Agents SDK. You need it when the agent is embedded in a product: your own interface, your own tools, your own rules.
Anthropic's advice from Building effective agents: start with the API directly and simple patterns. Frameworks speed up the start but hide prompts and responses, which makes debugging harder.
3. Tools
Without tools, the model can only answer with text. An agent's basic set is reading and editing files, search, a terminal, the web. Everything else connects through MCP: a task tracker, a database, a browser, error monitoring. More on the connected stack in the seventh post.
If you write tools yourself, the conclusions of Anthropic's Writing effective tools for agents are useful:
- Fewer, but better. More tools is not better. A single
schedule_eventis more useful than the combination oflist_users+list_events+create_event. - Clear names with prefixes, such as
asana_searchandjira_search, so the agent doesn't confuse similar tools. - Meaningful responses. Human-readable names instead of UUIDs, pagination and filters, clear error messages.
- A tool description is a prompt. Small edits to descriptions yield a noticeable gain in quality.
4. Context
Context is everything the model sees at the moment of decision: the system prompt, instructions, tool descriptions, history, call results. Anthropic calls working with it context engineering: finding the smallest set of tokens with the greatest value.
The main reason is context rot, the degradation of long context: the model has a limited “attention budget,” and every new token spends part of it. So:
- keep the instructions file short and specific;
- give the agent links and search tools, not the whole knowledge base at once. Anthropic calls this “just in time” context loading;
- for long tasks, use compaction, notes in external files, and subagents that return only the result.
If the agent answers from your documents, that is a RAG task: find the relevant fragments and insert them into the context at request time.
5. Checks
Without checks you don't know whether the agent works or only looks like it works. Two levels:
- Checks within a task: tests, build, linter that the agent runs itself (the verification loop).
- Checks of the agent as a whole: evals, a set of tasks with success criteria. According to Anthropic's recommendation, 20–50 tasks taken from real failures are enough. You can grade them with code (fast and objective), with a judge model against a rubric (flexible) or with a human (accurate but expensive). Be sure to read the transcripts: without that you won't know whether the checks themselves work correctly.
6. Permissions and sandbox
The principle of least privilege: the agent gets only the tools and access the task requires. OWASP lists excessive agency among the ten main risks of LLM applications and recommends limiting the set of extensions, giving them minimal permissions and requiring human approval for consequential actions.
In practice: deny rules for secrets and dangerous commands, a mode with confirmations for unfamiliar projects, a sandbox for shell commands, separate narrowly scoped tokens for MCP servers.
7. Budget and observability
Agents are expensive. In the article on its research system, Anthropic gives the numbers: an agent uses about 4 times more tokens than chat, a multi-agent system about 15 times. In their evaluation, the number of tokens used explained 80% of the variance in quality. The conclusion: a multi-agent setup is justified only when the value of the task covers the cost.
What to set up:
claude -p ... --output-format jsonreturnstotal_cost_usd; this is a client-side estimate and may differ from your bill;- in an interactive session,
/contextshows what the context window has been spent on; - session transcripts are stored locally, and they are worth reviewing.
Minimal starter set
- A Claude subscription or an API key.
- Claude Code, Codex or Cursor.
- A CLAUDE.md or AGENTS.md under 200 lines and one check command.
- Permission rules: allow for safe things, deny for secrets and
git push. - A dozen tasks on which you compare results.
- A spending limit and the habit of looking at what the agent did.
How to assemble this step by step is described in the second post.
Terms in this post
Practise it in
Sources
- Claude Code — How Claude Code works (агентный цикл и харнесс)
- Claude — Models overview
- Anthropic — Building effective agents
- Anthropic — Writing effective tools for agents
- Anthropic — Effective context engineering for AI agents
- Anthropic — Demystifying evals for AI agents
- Anthropic — How we built our multi-agent research system
- Claude Code — Run Claude Code programmatically
- OWASP — LLM06:2025 Excessive Agency
- Claude Agent SDK — overview
