Before You Build an AI Agent, Configure One
Start with one bounded job in a general-purpose agent, then add tools, tests, permissions, logs, and custom software only when the work justifies it.
Fabian Mösli Reading Preferences
Key Takeaways
- • Building your own agent usually means designing a job around an existing model. It rarely means training an LLM.
- • Start inside an approved general-purpose agent with static files, a written job contract, read-only tools, and human review.
- • Move to a purpose-built service only when the task has stable inputs, defined outputs, repeatable checks, narrow permissions, a failure path, and an owner.
In this guide
“We should build an agent” can mean three completely different projects.
One person means setting up a dedicated workspace in ChatGPT, Claude, or a coding agent. Another means connecting a few apps in a visual automation builder. An engineering team hears authentication, APIs, databases, logs, monitoring, testing, and somebody carrying the pager when it fails.
All three can be valid. They have wildly different costs and commitments.
I took the third route for my own assistant, on purpose. I rented a server, installed an open-source agent framework by hand, and spent a weekend on it because I wanted to understand agents from the inside rather than from a keynote. It worked, and it is still running. It also cost me several setups I broke so thoroughly that restoring the whole server from a backup was the only way back.
That was a good trade for learning. It is a bad trade for a job with a deadline, which is usually what people mean when they say they want to build an agent at work.
My default is to configure an existing general-purpose agent first. Give it one bounded job, static files, clear instructions, and no live write access. Run the job several times with a person watching. That teaches you where the real ambiguity and risk sit before you turn a rough idea into software.
An API integration will not rescue an unclear job or an unstable process. If the work depends on live state, begin with the narrowest read-only connection you can test under supervision.
This is part two of The Agent Field Guide. Part one explains how to distinguish chat, workflows, agents, and autonomy. Here I will turn that mental model into a practical build sequence.
What a general-purpose agent is
“General-purpose agent” has no settled industry definition. I use it for a broad, reusable agent that can handle many kinds of goals inside one environment.
A coding agent may read and edit almost anything in a software project, run terminal commands, browse documentation, and manage version-control changes. A work agent may use files, spreadsheets, email, calendars, and connected business apps. A research agent may search widely, analyse documents, and produce reports.
None is truly universal. Each is general-purpose within the workspace and tools it has been given.
The advantage is flexibility. I can ask the same coding agent to add a guide today, investigate a broken build tomorrow, and review a pull request next week. I do not need a separate application for each job.
The trade-off is management. I choose the goal, provide missing context, judge the result, and decide what happens next.
You can make that generalist much better without writing new software. Add project instructions, curated knowledge, narrowly approved connectors, and reusable procedures. I call the result a configured general-purpose agent: a generalist you have onboarded into your environment and way of working. You still assign different jobs and judge each result.
My project memory guide shows what those layers look like in one current product. The pattern applies well beyond it.
What makes a purpose-built agent different
A purpose-built agent accepts one known class of work under a standing operating contract. Its inputs, outputs, boundaries, failure path, and owner are designed once rather than reconstructed by the user for every task.
Think of a configured general-purpose agent as a well-equipped workplace for a person. A purpose-built agent is a repeatable job that somebody owns when it goes wrong.
| Configured general-purpose agent | Purpose-built agent |
|---|---|
| Accepts varied goals assigned task by task | Accepts one known class of work from a defined trigger |
| Offers a broad workspace and toolbox | Receives only the job-specific context, tools, and authority |
| The person selects or supplies inputs for the current task | Expected inputs are defined and checked before every run |
| The output follows the current request | The output or system change follows a stable contract |
| The person sets task-specific success criteria and judges the current result | The same outcome checks are built into every run |
| Approvals are chosen for the current job | Approval points are built into the service |
| Problems return to the current user | Recovery and escalation go to a named owner |
Memory, schedules, narrow tools, and automated checks do not settle the category. A general-purpose agent can have all four. Running the same task every Monday does not automatically turn it into an operational service.
The boundary is also independent of code. A no-code setup can be purpose-built. A thousand lines of agent framework can still produce a vague demo.
When purpose-built is worth the extra cost
Stay with the configured generalist while the process is changing, a knowledgeable person runs every case, or the output remains a draft. The general-purpose environment gives you flexibility and exposes assumptions cheaply.
A purpose-built service starts to make sense when several conditions hold:
- the job repeats often enough to justify maintenance;
- the expected inputs and outputs have stopped changing every week;
- several people need the same result;
- live systems or event triggers must be connected;
- the checks and approval boundaries can be stated clearly; and
- the organisation needs consistent logs, access controls, cost limits, and ownership.
Volume alone is not enough. Automating an unstable process at scale creates a faster source of exceptions. And “regulated” does not automatically mean custom code is safer. A sanctioned platform with mature identity, audit, and data controls may be the better foundation until a custom system can match them.
Choose your commitment level
There are three useful ways to build.
Configure a general-purpose agent
Add instructions, files, reusable procedures, and approved connectors inside a product you already use. This is the right starting point for personal work, internal drafts, research, and reversible tasks. You get most of the learning with the least machinery.
Build a no-code or low-code process
Add a form, schedule, or event trigger; let an agent handle the ambiguous middle; then return to fixed workflow steps and human approvals. This works when the process repeats and existing connectors cover the systems involved.
Build an agent application
This is a software product. It needs user identity and access controls, custom connections, task state, logs, testing, failure recovery, monitoring, and ongoing support. Choose it when the required control, integration, audit evidence, scale, or product differentiation justifies owning that software.
Start at the top. Move down only when a tested limitation forces the decision.
Write the job before you build the agent
Avoid broad ambitions such as “an assistant for marketing.” Pick one job you already understand.
I will use a weekly campaign review throughout this guide. The agent receives three approved sources, checks performance and spend, flags missing evidence, and drafts decisions for a person to approve.
Write the operating contract before touching a builder or software development kit (SDK):
Job: Prepare the weekly campaign review.
Trigger: I start it manually every Monday.
Inputs: Approved analytics export, campaign plan, and spend sheet.
Allowed tools: Read those three sources; calculate totals in a spreadsheet.
Deliverable: One-page review with results, anomalies, and proposed decisions.
Definition of done: Every figure matches a named source; gaps are flagged.
Ask for approval before: Updating the campaign tracker or sending anything.
Never: Guess a missing number or alter a source file.
On failure: Stop and list the missing data or failed tool.
Owner: Campaign lead.
This contract does more useful work than choosing an agent framework. It makes the vague parts visible: what starts the job, where facts come from, which actions are permitted, how success is checked, and who deals with failure.
If you cannot fill in those fields, you are not ready to automate the work. Run it manually and learn the process first.
Build it in six stages
1. Run the job in a general-purpose agent
Open a dedicated project or workspace in the AI tool your company already approves. Put the contract into its project instructions, add the three source files, and ask for a draft. Do not connect live accounts yet.
Run the same job beside it. Watch where the agent hesitates, guesses, asks the wrong question, or reaches for the wrong source. The first version may be nothing more than one project, one instruction file, and a few documents.
You are learning the real process before automating it.
2. Turn what worked into a reusable procedure
Save the instructions as a reusable skill, template, or standard operating procedure. Include examples of good output and the edge cases you found. Describe the available tools so clearly that the model can tell them apart.
Do not encode the process you wish existed. Encode the one that survived real cases.
This step is what makes an agent feel like yours instead of generic. For my own setup I wrote one procedure for adding a tool review to this site: the exact format, where the file goes, which fields matter, how my take should read. After that I stopped re-explaining it. It is the same move as onboarding a sharp new hire — capable on day one, clueless about how your shop does things until somebody writes it down.
3. Add the smallest useful tool set
Start with reading. Let the agent search documents, inspect records, and prepare a proposed action. Add write access only after the read-only version has earned it.
One narrow action called “create a draft campaign review” is safer and easier to test than permission to edit every record in a marketing system. Good tool design limits the mistakes the model can make.
4. Put gates in front of consequences
Require approval before external messages, payments, record changes, production deployments, deletions, or decisions about people.
The gate on my own website agent is one message long. She makes the change and builds a preview, then sends me a link. Nothing goes live until I reply that it looks good. That single step is why I am relaxed about letting her edit a public site from my phone, and it costs me about five seconds.
Keep rule-based checks in software or workflow rules. Required fields, amounts, permissions, and policy thresholds should not depend on an LLM remembering them at the right moment.
This is where code, automation, and AI take different jobs. Strong systems combine them instead of asking the model to improvise everything.
5. Test outcomes, not confidence
Collect real cases: normal examples, missing information, conflicting instructions, unusual formats, tool failures, and known past mistakes. Decide what must be true at the end of each run.
If an agent says it created a record, check that the record exists with the right fields. If it says the tests pass, run the tests. Anthropic’s agent-evaluation guide draws the same line between a convincing transcript and the actual state of the environment.
Keep every useful failure. Those cases become a regression set that you rerun whenever you change the model, instructions, knowledge, or tools.
6. Add unattended operation last
Schedules, background triggers, and automatic writes come after the supervised version works. Set a maximum number of steps, a time or spending limit, and a clear hand-off when the agent gets stuck.
You now need a record of every run: which version ran, what it read, which tools it called, what changed, what it cost, who approved it, and why it stopped. If several people rely on the service, somebody must own maintenance and incident response.
The safety bar rises with authority
An agent that drafts into a temporary file can waste your time. An agent that sends messages, changes production data, or handles personal information can harm other people.
At work, use only your sanctioned company setup and approved data sources. Connecting a personal AI account to company email, customer records, health data, or employee files is a security and data-protection problem, however convenient the workflow looks.
Approval is the starting point. Also confirm that the use is permitted, send only the minimum data the job needs, know how long the provider keeps it and which other vendors receive it, and decide who may read the prompts and run records. Those records can contain the same sensitive material as the original inputs.
My own approach here is boring on purpose, and I mean that as a recommendation. The agent runs on a cheap dedicated server that holds nothing else, so the worst outcome is a rebuild. It gets the access a given job needs and not the keys to everything else.
Then apply a few boring controls:
- Grant the minimum access needed for the job. Read-only first.
- Keep early runs in a sandbox, copy, draft folder, or test account.
- Require a person to approve irreversible or externally visible actions.
- Treat webpages, emails, and documents as untrusted input. A file can tell the agent to ignore its rules and upload a customer list. Enforce permissions and approval gates outside the content the agent reads.
- Keep backups and a tested way to undo changes.
- Verify the changed system rather than trusting the agent’s claim that it succeeded.
- Record enough of the run to reconstruct a failure.
OpenAI’s practical guide to building agents recommends rating tools through factors including read versus write access, reversibility, account permissions, and financial impact. That is more useful than one global permission switch.
For systems that influence legal, employment, credit, health, or other high-impact decisions, a promising prototype is nowhere near production. Bring in engineers, security and privacy specialists, domain owners, and lawyers before the system gains real authority.
Your first version should feel almost too small
Choose one weekly job. Write the contract. Put three static source files into a dedicated workspace in the approved AI product you already have. Ask for a draft. Keep live accounts disconnected.
Run that version five times. Save the failures. Tighten the instructions and checks. Only then decide whether the next limitation calls for a connector, a workflow, or custom software.
That is how you build your own agent without starting a software project by accident.
If the terms “agent,” “workflow,” and “autonomy” still feel slippery, start with part one: Is It Really an AI Agent?. For a concrete general-purpose coding agent, read Claude Code: When AI Stops Talking and Starts Doing. To practise delegation and permission-setting without risking a real system, play The Company Simulator.
Published: 2026-09-15
Last updated: 2026-09-15