Understand loops, autonomy, handoffs, and the point where an agent should stop.
An AI agent is a system in which a language model decides what to do next, takes an action through a tool, reads the result, and repeats until a goal is met or a limit is hit. That one change, letting the model choose the next step instead of a programmer hard-coding it, is what makes agents useful for open-ended work such as debugging, research or operations. It is also what makes them expensive, slow and hard to predict: every extra step is another chance to be wrong, and small per-step error rates compound over a long run.
The useful mental model is a loop wrapped in a harness. The loop is simple: the model reads its context, emits either a tool call or a final answer, the runtime executes the call, and the result goes back into context. Almost all of the engineering lives outside the model: which tools exist and with what permissions, how many steps and how much money a run may spend, where a person must approve, how state is saved so a crash does not lose an hour of work, and how the system decides the task is really done rather than taking the model's word for it.
The questions build in that order. The Basic questions define agents, separate them from chatbots, explain the loop, and introduce the control ideas: decomposition, human approval, bounded autonomy, and the planner and reviewer roles. The Advanced questions cover the architecture decisions that separate demos from production: when several agents beat one, when they do not, how to stop runaway loops, how explicit state machines and checkpoints make runs safe and resumable, how to define done, how handoffs work, and the failure modes that show up once real users arrive.
An AI agent is a system where a language model repeatedly chooses an action, executes it through a tool, observes the result, and decides the next step until it reaches a goal or a stopping limit.
A plain LLM call is one shot: text in, text out. An agent puts that call inside a loop and gives the model tools, which are functions the surrounding program can run on its behalf: search a database, read a file, call an API, run tests. On each turn the model looks at the goal and everything that has happened so far, then either requests a tool call or returns a final answer. The program runs the tool, appends the result to the context, and calls the model again.
The defining property is that the model controls the path. In a fixed pipeline the programmer decides that step two always follows step one. In an agent the model decides whether to search again, try a different file, or stop. That is what lets an agent handle tasks where the steps depend on what it discovers, such as tracking down why a test fails: the next useful action depends on what the last stack trace said.
An agent therefore has four parts: a model that reasons and chooses, tools that act on the world, context that carries the goal, instructions and observations, and a harness (the runtime code around the model) that executes tools, enforces limits and stores state. Most quality and safety problems are solved in the harness and tools, not in the model.
The trade-off is control. Agents take more steps, cost more tokens, and vary from run to run. Each step can fail, and those failures compound. For a task with known steps, a deterministic workflow with an LLM call at specific points is cheaper, faster and easier to test. Use an agent when the path genuinely cannot be written down ahead of time.
A chatbot produces a reply for a person to read and act on; an agent takes several actions itself through tools, observes results and continues toward a goal, so its errors become actions rather than just wrong text.
A chatbot is a conversational interface: a person writes, the model replies, and the person decides what to do next. Even when a chatbot can call a tool, such as looking up an order status, it usually does one lookup and answers. The human is the loop. They read the reply, judge it, and type the next message.
An agent moves the loop inside the system. Given a goal like "reschedule my meeting with Priya to a time we are both free next week", it checks both calendars, finds a slot, sends the invite and confirms, without a person approving each step. The model decides how many steps to take and when it is finished.
The real difference is who carries the consequences of a mistake. A chatbot's error is a bad sentence that a person can ignore. An agent's error is a sent email, a deleted record or a refund issued. That shifts the engineering focus from answer quality to action safety: permissions, approvals, idempotent tools, audit logs and budgets.
The line is a spectrum, not a binary. Many products sit in between: a chat interface that can take one or two confirmed actions, or an agent that pauses to ask a question. Choose the point on the spectrum by the cost of an unsupervised error, not by how impressive the demo looks.
An agent loop is the repeated cycle in which the model reads context, chooses a tool call or a final answer, the harness executes the call, and the result is added to context before the next turn.
The loop is the engine of every agent and is usually only a few dozen lines of code. The harness sends the model the goal, instructions, tool definitions and the transcript so far. The model replies with either tool calls (a tool name plus JSON arguments) or a final answer. If it asked for tools, the harness runs them, appends each result to the transcript as an observation, and calls the model again. The loop ends when the model returns a final answer or the harness hits a limit.
This is the pattern popularised by the 2022 paper ReAct: Synergizing Reasoning and Acting in Language Models: interleave reasoning with actions so each decision can use fresh evidence. Modern provider APIs build it in through native tool calling, so you rarely parse actions out of free text any more.
The model does not execute anything. It only proposes. The harness is where you decide whether a call is allowed, validate its arguments, run it with the right credentials, catch errors and turn them into readable observations, and count steps and cost. That split is what makes agents controllable: every action passes through code you own.
Two loop properties matter most in practice. First, context grows every turn, so long runs get slower, more expensive and less accurate unless you trim or summarise old observations. Second, the loop needs exits other than the model saying it is done: a step cap, a cost cap, a wall-clock timeout and a check for repeated identical calls.
Task decomposition is breaking a large goal into smaller subtasks that each have a clear input, output and check, so each piece fits in context, can be verified, and can fail without sinking the whole run.
Models do well on bounded, well-specified tasks and drift on long, vague ones. Decomposition turns "migrate the billing service to the new payments API" into pieces like: list every call site, map old fields to new ones, change one module at a time, run that module's tests. Each piece has a smaller context, a narrower set of tools and a check that tells you whether it worked.
Decomposition helps for three reasons. Context: each subtask needs only the files and facts relevant to it, which keeps attention focused. Verification: a small step with a concrete output (a list, a diff, a passing test) can be checked by code before the next step builds on it. Recovery: when step 4 of 9 fails you retry step 4, not the whole job.
There are two ways to do it. Up-front decomposition produces a plan before any action, which suits tasks whose structure is predictable. Incremental decomposition decides the next subtask after each result, which suits exploratory work like debugging where you cannot know step 3 until you see step 2. Most good agents mix them: a rough plan that is revised as evidence arrives.
The cost is coordination. Every boundary between subtasks is a place where context is lost and has to be passed explicitly, and too fine a split means many model calls for little work. A useful size is a subtask a competent person could do in a few minutes to an hour, with an output that can be checked.
Human-in-the-loop means a person participates at defined points in an agent's run, approving, correcting or supplying information, so that high-stakes or ambiguous decisions are not made by the model alone.
In a human-in-the-loop (HITL) design, the agent pauses at specific points and waits for a person. The common forms are approval ("issue this 4,200 rupee refund?"), review and edit (a person corrects a drafted reply before it is sent), clarification (the agent asks which of two customers named Sharma was meant), and escalation (the agent hands the case over because it hit a limit or a policy).
HITL is worth its friction where an error is costly or irreversible: money movement, external communication, data deletion, production changes, legal or medical judgement. It is also how you build trust in a new agent. Start with approval on every write, measure how often people change or reject the proposal, and relax the gate for action types that are approved unchanged almost every time.
The engineering detail that matters is that the pause must be durable. A person may take ten minutes or two days to respond. The agent's state, including the exact proposed action and its arguments, must be saved so the run can resume after approval without re-running the model, which might propose something different the second time. The approval should bind to that exact action: approving "refund 4,200" must not let a later step refund 42,000.
The failure mode of HITL is rubber-stamping. If a reviewer sees fifty requests an hour with no context, they click approve on all of them, and the gate provides the appearance of safety without the substance. Show the reviewer what will happen, why, and what evidence supports it, and keep gates for decisions that actually need judgement.
Bounded autonomy lets an agent decide freely inside explicit limits enforced by the harness: which tools and data it can touch, how much it can spend, how long it can run, and which actions need a person.
An agent with no limits is a liability, and an agent that needs approval for everything is a slow form. Bounded autonomy is the middle: the agent chooses its own steps, but only inside a box defined by code. The box has several walls. Scope: which tools, which repositories, tables or customers. Effect: read-only, reversible writes, or irreversible actions. Budget: maximum steps, tokens, money and wall-clock time. Escalation: which situations must stop and ask.
The key word is enforced. Writing "never deploy to production" in the system prompt is a request, not a boundary. A prompt injection in a web page, a confused plan or a simple misreading can override it. Real boundaries live where the model cannot argue with them: the tool is not in the agent's toolset, the credential it runs with cannot write to production, the database role is read-only, the sandbox has no network.
Bounds should be set by blast radius, meaning the worst plausible outcome of a mistake. An agent triaging support tickets can label and route freely because a wrong label is cheap to fix. The same agent should not close tickets or issue credits without a check. Coding agents commonly get free rein inside a sandboxed branch but cannot merge, push to main or touch secrets.
Bounds are not static. Start narrow, log every action, and widen a bound only when the data shows the agent handles that class of action well. Narrowing must be quick too: a feature flag that removes a tool is more useful in an incident than a prompt edit.
A planner agent turns a goal into an explicit, ordered set of steps with expected outputs and checks, usually without acting itself, so execution can be reviewed, parallelised and tracked against the plan.
A planner separates deciding what to do from doing it. Given a goal and some context, it produces a plan: a list of steps, each with a description, dependencies, the expected output and how to verify it. One or more executors then carry out the steps. This is often called plan-and-execute, as opposed to a single agent deciding one step at a time.
Planning up front has real benefits. A person can review the plan before any side effect happens, which is a cheap and effective approval point: approving a ten-line migration plan is easier than approving forty individual tool calls. Steps without dependencies can run in parallel. The plan becomes a progress tracker, so a resumed run knows which steps are done. And the plan gives the reviewer something to check the result against.
The weakness is that plans are made with the least information the run will ever have. The first real observation often invalidates part of the plan: the file is not where expected, the API returns a different shape. A planner that never revisits its plan produces an executor that faithfully does the wrong thing. Good designs replan at checkpoints: after a step fails, after a surprising result, or every few steps.
Make the plan structured, not prose. A JSON list of steps with ids, dependencies and acceptance checks can be validated, displayed, stored and diffed. Keep plans short. A plan with forty steps for a task that needs six is usually padding, and each extra step is another chance to drift.
A reviewer agent checks another agent's output against explicit requirements and evidence, then approves it or returns specific defects, giving the system a second look that catches errors the producer is blind to.
A reviewer (also called a critic or verifier) takes an artefact, such as a code diff, a drafted email or a research summary, plus the requirements it should meet, and returns a verdict with reasons. If it finds problems, the producer revises and the reviewer checks again, up to a fixed number of rounds.
Reviewers help because producing and checking are different tasks. A producer that just wrote an answer tends to read it charitably, a self-consistency bias that is easy to observe. A reviewer with a fresh context, a narrower brief ("find unvalidated inputs") and sometimes a different model or prompt catches errors the producer glosses over. It works best when the reviewer can check against something concrete: the spec, the source documents, the test output.
Reviewers have known weaknesses. They share many blind spots with the producer, especially when the same model is used. They can be talked into approval by confident prose. They tend to over-report small style issues while missing a substantive bug. And an LLM reviewer is still a model, so its approval is evidence, not proof. Wherever a deterministic check exists (tests, a schema validator, a linter, a citation matcher), run it first and let the model review what code cannot check.
Design the reviewer's output for action. Ask for a verdict, a list of defects each tied to a location and a requirement, and a severity. Cap revision rounds at two or three; if the work still fails, escalate to a person rather than letting producer and reviewer argue indefinitely.
Multiple agents help when work splits into independent parts that can run in parallel, when subtasks need isolated context or permissions, or when a separate checker adds real verification; otherwise coordination costs usually outweigh the gains.
A multi-agent system uses several model-driven loops that each own part of the work, usually coordinated by an orchestrator agent or by code. The question is never whether multiple agents are more sophisticated. It is whether splitting the work produces better results per unit of cost and latency than one agent with good tools.
There are four situations where the split pays. Parallel breadth: a research task that needs twenty independent sources read can fan out to sub-agents that each read a few and return a condensed summary, cutting wall-clock time and keeping each context small. Context isolation: a sub-agent can read 200,000 tokens of logs and return 300 words, so the orchestrator's context stays clean. This is often the biggest benefit, and it is really a context-management technique. Permission isolation: an agent that reads untrusted web content should not hold the credential that sends emails; separating them limits what a prompt injection can reach. Independent verification: a reviewer with a fresh context and a different brief catches errors the producer misses.
What these share is that the parts are loosely coupled. Each sub-agent gets a clear, self-contained brief and returns a compact, structured result. Tasks where every step depends on detailed shared state, such as editing one tightly coupled codebase, divide badly, because each handoff loses context that the next agent needs.
Multi-agent systems cost more. Every agent re-reads its own instructions and context, so token use can be several times that of a single agent on the same task. Debugging spans multiple traces. Failures compound across handoffs. Build them when you can point to the specific benefit (time saved by parallelism, a context that would otherwise overflow, a permission boundary) and measure it against a single-agent baseline on the same eval set.
Multi-agent is unnecessary when one agent with good tools fits the task in its context, when steps are sequential and share state, or when the split exists only to mirror an org chart; then it adds cost, latency and failure points.
The default should be one agent. A single loop with well-designed tools, a clean context and a verification step solves most tasks, and it is far easier to debug: one trace, one context, one place where the decision was made. Teams often build multi-agent systems because the architecture diagram looks capable, then discover that the agents spend most of their tokens briefing each other.
Multi-agent is unnecessary in several recognisable cases. The task fits one context: summarising a 20-page document, answering a support question, or fixing a bug in one module. The steps are sequential and share state: each step needs the full detail of the previous one, so splitting means repeatedly serialising and re-reading the same information. The roles are cosmetic: a "researcher" agent that calls search and a "writer" agent that writes are just two prompts that a single agent could switch between. A tool would do: if the "specialist agent" always performs the same fixed operation, make it a deterministic function the main agent calls.
The costs are concrete. Error compounding: if each handoff has a 90 percent chance of passing the needed context intact, three handoffs succeed only about 73 percent of the time. Token overhead: every agent re-reads its system prompt and task brief, and conversation between agents is pure overhead. Coordination failures: agents duplicate work, contradict each other or wait on each other. Debuggability: a failure might originate in any agent's context, so root-causing takes much longer.
A useful test: describe what each agent knows that the others do not, and why that separation helps. If the answer is "nothing, it just has a different role prompt", collapse them. If you need different instructions for different phases, swap the system prompt or tool set within one agent instead.
Prevent runaway loops with hard limits enforced by the harness (steps, tokens, cost, wall-clock time), plus detection of repeated actions and lack of progress, and a defined exit that stops cleanly or escalates to a person.
Agents loop for predictable reasons. A tool keeps failing and the model keeps retrying with the same arguments. The model cannot find what it needs and keeps searching with small variations. Two agents hand a task back and forth. A test fails, the model makes a change, the test still fails, and it reverts and tries again. The model is not aware it is stuck; from inside the loop, each next step looks reasonable.
The first layer is hard budgets that the model cannot override: a maximum number of steps, total tokens, total spend and wall-clock time per run. These are backstops, not tuning knobs. Set them from data: if 95 percent of successful runs finish in 12 steps, a cap of 30 catches loops without cutting off legitimate long tasks.
The second layer is stall detection, which catches loops long before the budget runs out. Track a fingerprint of each tool call (name plus normalised arguments) and stop or intervene when the same call repeats several times. Count consecutive errors from the same tool. Track a progress signal where one exists, such as the number of failing tests or plan steps completed, and stop if it has not improved in N steps. For multi-agent systems, cap the number of handoffs per task.
The third layer is what happens at the limit. Do not silently fail. Inject a message telling the model it is stuck and must try a different approach or summarise what it learned, then give it one or two final steps. If that fails, end the run with a structured status ("stalled: search_orders returned no results 5 times") and route to a person with the trace. A run that stops with a clear reason is a success of the harness, not a failure.
Tool design prevents many loops in the first place. Errors that say what went wrong and what to try ("no customer with that email; try search_customers with a partial name") help the model change course, while a bare "error 400" invites the same call again.
An agent state machine defines the explicit states a run can be in and the allowed transitions between them, so the harness, not the model, controls the lifecycle, and illegal jumps like sending before approval become impossible.
A free-running agent has one implicit state: "the model is deciding". That makes it hard to answer basic questions: has this draft been approved? Is this run waiting on a person or stuck? Can it be resumed? A state machine makes the lifecycle explicit. A run is always in exactly one named state (for example drafting, in_review, awaiting_approval, sending, done, failed), and only listed transitions are allowed.
The value is that the harness enforces the transitions. The model can do whatever it likes inside the drafting state, using whatever tools that state allows, but it cannot move the run to sending: that transition requires an approval record. Each state can have its own tools, prompt and budget, so the drafting phase has search and write tools, while the sending phase has only the send tool with the approved content. This gives you the flexibility of an agent inside each state and the predictability of a workflow between them.
State machines also make operations tractable. Persist the current state and transition history to a database and you can show progress in a UI, resume from the last state after a crash, query how many runs are waiting for approval, and measure where runs fail. Transition events form a clean audit trail: who or what moved the run, when, and on what evidence.
Do not over-model. If you need twenty states with complex conditions, the task is probably a workflow with LLM steps, and a workflow engine may serve better than a hand-rolled machine. Graph-based agent frameworks express the same idea as nodes and edges; whichever you use, the principle is that lifecycle and permissions belong to code.
Agents resume by persisting durable state after every step (transcript or summary, plan progress, tool results and side effects already done) and restarting from the last checkpoint, using idempotent tools so replays never repeat an action.
Long runs get interrupted: a deploy restarts the worker, a provider times out, a person takes a day to approve, a budget pauses the run. Without persistence, the only option is to start over, which wastes money and can repeat side effects. Resumability means a new process can pick up the run and continue as if nothing happened.
What to persist after each step: the run's identity and state (goal, current state, owner), the plan and which steps are done, the transcript or a compacted summary of it, every tool call with its arguments and result, and a side-effect ledger recording which external actions completed and their ids (the refund id, the commit hash, the sent message id). Store it in a durable database, not in process memory or the model's context.
The hard problem is the crash between action and record: the email was sent, but the process died before saving that fact. On resume, the agent may send it again. The fix is idempotency. Generate an idempotency key for each side-effecting call from the run id and step number, pass it to the external API where supported, and check the ledger before executing. Many payment and messaging APIs accept such keys and will not repeat an operation that already succeeded under the same key. Where an API does not, query its state ("does a message with this key exist?") before retrying.
On resume, do not blindly replay the full transcript. Rebuild a clean context: the goal, the plan with completed steps marked, a summary of what was learned, and the last few observations. This is cheaper and often better than the original context, which may have been bloated with stale tool output. Record the resume as an event in the trace so the run's history stays honest.
Durable-execution engines and workflow orchestrators provide much of this machinery (checkpointing, retries, timers) and are worth considering once runs routinely span minutes to days.
A task is done when observable acceptance criteria defined before the run pass checks run by code or a person, such as tests passing or records matching, not when the agent says it is finished.
Agents are prone to premature completion: declaring success when the work is partial, the tests were never run, or the change was made in the wrong place. The model's final message is a claim, produced by the same process that did the work. Treating it as proof is the most common way agents ship broken results.
The fix is to define acceptance criteria before the run and check them independently afterwards. Good criteria are observable and specific: the named test file passes, the full suite still passes, the API returns 200 for these five requests, the CRM record shows the new address, the report's totals match the source table within rounding. Vague criteria like "the code is clean" or "the summary is good" cannot be checked and should be turned into something that can.
Checks come in a hierarchy of trust. Deterministic checks (tests, schema validation, a database query, a diff of expected files) are strongest. Evidence checks compare the output to sources, such as whether every claim in a summary has a citation that supports it. Model judges grade against a rubric where code cannot, and must be calibrated. Human review is the final layer for high-stakes or subjective work. Use the cheapest strong check available, and combine them.
The harness should run the checks itself, not ask the agent to report results. If verification fails, feed the specific failure back to the agent for another attempt within budget, or end the run with a status of "incomplete" and the evidence. A clear incomplete status is far more useful than a false "done".
Also check for collateral damage. The criteria for "done" include what must not have changed: no unrelated files edited, no tests deleted or skipped, no extra records touched. Agents under pressure to pass a check sometimes weaken the check itself, so guard the tests and fixtures they could modify.
A good handoff passes a self-contained, structured brief (goal, what is done, evidence, open questions and constraints) rather than a raw transcript, and transfers ownership explicitly so exactly one agent or person is responsible next.
A handoff happens whenever responsibility moves: an orchestrator delegates to a sub-agent, a sub-agent returns results, a triage agent routes to a billing specialist, or an agent escalates to a human. Every handoff is a lossy compression. The receiver does not share the sender's context, so anything not written into the handoff is gone. Most multi-agent failures are handoff failures: the sub-agent did not know the constraint, or the human had to re-ask the customer everything.
A good handoff is a brief, not a transcript. It states the goal and why it matters, what has been done and the evidence (ids, links to records, test results), what is unknown or was tried and failed, the constraints (budget, permissions, deadlines, things not to touch) and the expected output format. Passing the full transcript seems safer but usually is not: it is expensive, buries the important facts, and can carry untrusted content or injected instructions into a context that has more privileges.
Ownership must be explicit. At any moment one agent or one person owns the task. Record handoffs as state transitions with an owner field, so you never get two agents acting on the same ticket or a ticket nobody holds. For escalations to humans, put the case in the queue people already work from, with the brief at the top and the trace one click away, and set a timeout with a defined fallback if nobody picks it up.
Handoffs back to the orchestrator need the same discipline. Ask sub-agents for a structured result with a status (complete, partial, blocked), the findings, the confidence and the evidence, capped in length. The orchestrator should treat this as data to verify, not instructions to follow, particularly if the sub-agent read untrusted content.
Production agents fail mostly through compounding step errors, bad tool design, context overload, premature completion and unsafe actions on untrusted input; you catch these with traces, trajectory evals on real failures and limits enforced by the harness.
Agent demos run a handful of friendly tasks. Production runs thousands of messy ones, and the failure modes are remarkably consistent across teams. Knowing them in advance lets you design against them rather than discover them from incidents.
Compounding errors come first. If each step succeeds 95 percent of the time, a 20-step run finishes cleanly only about 36 percent of the time (0.95 to the power 20). Short runs, verification after risky steps and the ability to recover from errors matter more than squeezing another point out of the model. Tool problems are next: vague descriptions, overlapping tools, huge unfiltered outputs and unhelpful error messages cause wrong calls and loops. Many "model" failures disappear when a tool is renamed, split or made to return a useful error.
Context overload shows up in long runs: as stale observations pile up, the model loses track of instructions and repeats work. Premature completion is the agent claiming success without verifying. Unsafe action on untrusted input is the security failure: a web page, email or document contains instructions the agent follows, a risk that grows with every tool the agent holds. Cost and latency blowouts come from loops, oversized contexts and multi-agent chatter.
Catching them needs three things. Tracing: every model call, tool call, argument, result, token count and decision recorded per run, so a failure can be replayed. Trajectory evals: a regression set built from real failed runs, checking required and forbidden tools, order and budgets as well as final answers, run on every prompt, tool or model change. Production monitors: success rate by task type, steps and cost per run, stop reasons, approval and override rates, and alerts on spikes. Review a sample of traces by hand every week; the patterns there feed the eval set.