Anthropic CCA-F: Agent Loops and Orchestration

An agent loop sounds simple: observe the situation, decide what to do, take an action, inspect the result, and repeat. The difficult part is deciding what state persists, which actions are available, when another model or agent should be involved, how failures are handled, and what evidence is strong enough to stop the loop. Those are architecture choices.

Agentic architecture and orchestration are major themes for CCA-F. Anthropic’s certification materials place agentic design alongside tool/MCP integration, prompt engineering, structured output, and Claude Code workflows. That means loops should be studied as complete systems rather than as a prompting pattern.

The fastest way to build intuition is to implement one loop with a small number of tools and an explicit state object. Then add uncertainty, tool failure, long context, a second agent, and a high-impact action. Each complication exposes a different control problem.

A loop needs a state model

The agent has to know what has already happened. Without state, it may repeat the same search, lose a decision, or attempt an action twice. State can include the goal, completed steps, unresolved questions, tool results, approvals, cost, and the reason the loop should continue.

Do not confuse state with raw conversation history. A transcript contains everything that was said; state contains what the workflow needs next. Summarizing completed work into explicit fields can reduce context and make the loop easier to inspect.

This distinction is central to AI agent behavior. The model generates the next decision from the information presented to it. Better state representation produces better decisions than simply appending more and more transcript.

Stopping conditions should be designed before autonomy

Every loop needs a completion rule. The goal may be satisfied, a maximum number of steps may be reached, the evidence may be insufficient, a cost limit may be hit, a tool may fail repeatedly, or a human may need to decide. Without these rules, the model can continue because continuation is always technically possible.

A good stop condition is observable. “Stop when the task seems done” is weak. “Stop when all required fields are populated and validation passes” is stronger. “Stop after three failed attempts and request human review” is stronger still.

Budgets are also control mechanisms. Cap tool calls, token use, elapsed time, and retries where appropriate. A production agent should be able to fail predictably instead of consuming resources indefinitely.

Orchestration can live outside the model

Not every step in an agentic workflow should be chosen by the model. Deterministic code is better for rules that must always execute, such as validating a schema, enforcing a spending limit, checking permissions, or committing a transaction. The model should make decisions where interpretation is genuinely required.

The agentic shift does not eliminate ordinary software engineering. It changes where probabilistic reasoning adds value. A mature architecture surrounds that reasoning with deterministic boundaries that make outcomes safer and easier to test.

This is especially important when a loop includes retries. The model can decide that another attempt is useful, but code can enforce that the attempt count has not exceeded policy and that the action remains safe to repeat.

Multi-agent systems need clear ownership

Adding another agent can provide specialization, separate permissions, or independent context, but it also adds communication and coordination overhead. The design should explain why one agent cannot perform the job safely and efficiently.

A useful pattern gives each agent a distinct domain and contract. A research agent may gather evidence without write permission. A planning agent may combine evidence into a proposed action. An execution agent may have access to a narrow set of tools. A supervising component can decide when work moves between them.

Concepts from agent-to-agent protocols reinforce the need for explicit communication. Messages should contain enough state to continue the task without requiring every agent to share the full transcript and every permission.

Tool routing is part of orchestration

An agent loop often fails because it chooses the wrong tool, not because the language model cannot reason. Similar tool names, vague descriptions, oversized catalogs, and noisy outputs all increase selection difficulty.

Use descriptive names and schemas, load only relevant tools, and return concise results. Model Context Protocol can standardize access to external tools and resources, but the architecture still decides which capabilities are exposed in each stage of the loop.

For complex workflows, tool discovery may itself become a step. The agent can search a large capability catalog and load only the handful of tools needed for the current goal. This reduces context pressure and improves routing accuracy when the integration surface is large.

Approvals create a deliberate break in the loop

High-impact actions often deserve human approval. The agent can prepare the proposed action, summarize evidence, and pause. A person can approve, reject, or modify the request before execution continues.

The approval object should be tied to exact parameters. If the agent proposes a $500 refund and the amount changes to $5,000, the old approval should not remain valid. If the data used to justify the action changes materially, the workflow may need new review.

This pattern gives the model autonomy in low-risk reasoning while preserving human authority over consequential actions. It also creates an audit trail that explains where responsibility changed hands.

Evaluation should inspect trajectories

Final-answer evaluation is not enough for agents. Two loops can produce the same answer while one uses five unnecessary tools, leaks data into context, or retries a dangerous action. Evaluate the trajectory: steps taken, tools chosen, arguments used, evidence gathered, cost incurred, and stop condition.

The general discipline of foundation model evaluation becomes more complex in agent systems because quality is partly behavioral. A good test suite includes tasks where the correct outcome is to ask for clarification, refuse an action, escalate, or stop.

Capture traces from successful and failed runs. Cluster common failure patterns. If the same unnecessary step appears repeatedly, fix the orchestration or tool surface rather than adding another sentence to the prompt.

Long loops need context hygiene

Every tool result, intermediate plan, and conversational turn can accumulate in context. After enough steps, the loop may spend more tokens rereading history than solving the current problem. More importantly, obsolete information can continue influencing decisions after the workflow has moved on.

Summarize completed stages into compact state, preserve only evidence needed for future decisions, and discard temporary artifacts when they no longer matter. If a tool returns a thousand records but the next step needs only three totals, keep the totals. If an approval closes a decision, store the approved parameters rather than the whole debate.

Context hygiene also makes debugging easier. A trace that separates active state from historical transcript lets an engineer see exactly what information shaped the next action.

Handoffs create trust boundaries between agents

When one agent hands work to another, define what authority crosses the boundary. A research agent may pass evidence but not credentials. A planning agent may pass a proposed action but not execute it. An execution agent may receive validated parameters and a signed approval object rather than the entire upstream conversation.

This limits accidental privilege propagation. It also lets each agent operate with a smaller context and tool set. The receiving agent needs enough information to do its job, not every thought produced earlier in the workflow.

Handoffs should be testable contracts. Define required fields, provenance, confidence or validation status where useful, and the error behavior when the payload is incomplete. Multi-agent orchestration becomes more reliable when agent boundaries look like well-designed service boundaries.

Build loops that become simpler under inspection

For CCA-F preparation, build an agent that must research a question, choose one of three tools, produce structured output, and request approval before a write action. Store explicit state. Set a maximum step count. Create a retry rule for one tool and a no-retry rule for another.

Then break it. Return conflicting evidence. Hide the correct tool among similarly named tools. Make the second agent unavailable. Change an approval parameter. Fill the context with irrelevant history. Observe whether the loop becomes confused or whether the architecture keeps the task bounded.

Cost is another useful signal inside the loop. A system can record token use, tool charges, elapsed time, and the number of planning iterations as part of state. If the expected value of another step is low, the architecture can stop or escalate instead of continuing automatically. This makes resource limits part of decision quality rather than a separate billing concern.

Reliability also improves when the loop distinguishes “unknown” from “failed.” Unknown means more evidence may help. Failed means the system attempted an action and could not complete it. Those states should lead to different next steps, and collapsing them into one generic retry can create repeated errors or unnecessary calls.

The goal is not maximum autonomy. It is controlled autonomy that can be explained. A well-designed agent loop has clear state, narrow tools, explicit stopping rules, visible handoffs, bounded retries, and enough telemetry to reconstruct what happened after the fact.

img