Agent Design Principles

Reference intermediate

Design rules for building agents you can ship. An agent here means an LLM in a loop: it observes, picks an action (usually a tool call), runs it, observes the result, and repeats until it hits a goal or a stop condition. Each rule is a default; deviate only when you can say why.

The Principles at a Glance

# Principle One-line rule
1 Observable loop Log every think → act → observe step as structured data.
2 Narrow tools One tool, one job, with a precise description and schema.
3 Fail safe Errors return recoverable data, not exceptions that crash the loop.
4 Bound resources Hard caps on steps, tokens, time, and money.
5 Human in the loop Gate irreversible or high-blast-radius actions behind approval.
6 Explicit state Store goal, plan, and progress as data, not only in the transcript.
7 Least privilege The agent gets the smallest capability set the task requires.

1. Keep the Loop Observable

You cannot debug what you cannot see. Emit one structured (JSON) record per iteration with the model's reasoning, the chosen tool, its inputs, the raw result, and running token/cost counters. The trace is your primary debug artifact.

Log this Why it matters
Reasoning text Reveals why a bad action was chosen
Tool name + args Reproduce any step in isolation
Raw tool result Separate model error from tool error
Cumulative tokens/cost Catch runaway loops early
Stop reason Finished, capped, or errored?

2. Make Tools Narrow and Well-Described

The model selects and fills tools almost entirely from their description and schema. Vague tools cause wrong calls. Prefer many small tools over one with a mode flag, and write each description as if it is the only docs the model gets — it is.

search_orders = {
    "name": "search_orders",
    "description": (
        "Find orders for a single customer by email. Returns up to 20 "
        "most recent orders, newest first. Use before refunds to confirm "
        "an order exists. Does NOT search by order ID — use get_order."
    ),
    "input_schema": {
        "type": "object",
        "properties": {
            "email": {"type": "string", "description": "Customer email, lowercase"},
            "since": {"type": "string", "description": "ISO 8601 date; omit for all time"}
        },
        "required": ["email"],
    },
}
Smell Fix
Tool named do_stuff / helper Name for the action: cancel_subscription
A mode param that switches behaviour Split into two tools
Description omits return shape State what comes back and its limits
Free-form string for an enum Use "enum": [...] in the schema
Side effect unmentioned Say "writes to DB" / "sends email"

3. Fail Safe

A crashing tool kills the agent; a tool that returns an error string lets the model adapt. Catch exceptions at the tool boundary and hand back a message it can act on. Validate inputs against the schema first.

def err(msg):
    return {"type": "tool_result", "is_error": True, "content": msg}

def execute_tool(name, args):
    if name not in TOOLS:
        return err(f"Unknown tool '{name}'. Available: {list(TOOLS)}")
    try:
        return {"type": "tool_result", "content": TOOLS[name](**args)}
    except Exception as e:
        return err(f"{type(e).__name__}: {e}")

Rules of thumb: never silently swallow an error, never auto-retry a destructive call, and make messages actionable ("file not found: report.csv" beats "IOError"). Hard-stop only on unrecoverable conditions (auth failure, budget exhausted).

4. Bound Every Resource

An unbounded loop is a runaway bill or an infinite cycle. Enforce limits in code, not the prompt — the model will not reliably count.

Limit Default start Enforced by
Max steps 15–25 per task Loop counter
Token budget Cap per task Sum from API usage
Wall-clock 60–120 s Deadline timestamp
Tool calls Per-tool cap Counter dict
Spend Per-task ceiling Cost accumulator
import time

def run(client, goal, max_steps=20, deadline_s=120):
    messages = [{"role": "user", "content": goal}]
    start = time.monotonic()
    for step in range(max_steps):
        if time.monotonic() - start > deadline_s:
            return {"status": "timeout", "step": step}
        resp = call_model(client, messages)   # your API wrapper
        if resp.stop_reason != "tool_use":
            return {"status": "done", "result": resp}
        messages.append({"role": "assistant", "content": resp.content})
        results = [execute_tool(b.name, b.input)
                   for b in resp.content if b.type == "tool_use"]
        messages.append({"role": "user", "content": results})
    return {"status": "max_steps"}

5. Keep a Human in the Loop for Risky Actions

Classify each tool by blast radius. Read-only and easily reversible actions run autonomously; irreversible or high-impact ones require approval. The gate lives outside the model — it can request, but only a person (or your policy) authorizes.

Tier Examples Policy
Auto read file, search, read-only query Run without approval
Confirm write file, send email, create ticket Show diff/preview, then approve
Block delete records, deploy, spend money, rm -rf Require human sign-off
RISKY = {"delete_records", "deploy", "issue_refund", "send_email"}

def gate(tool, args, approve_fn):
    # approve_fn blocks for a human/policy decision and returns a bool
    if tool in RISKY and not approve_fn(tool, args):
        return {"is_error": True, "content": "Action denied by approver."}
    return execute_tool(tool, args)

6. Prefer Explicit State

The transcript is lossy and gets compacted. Keep the goal, plan, completed sub-tasks, and key facts in a structured object you control, and re-inject the relevant slice each turn rather than trusting the model to remember.

state = {
    "goal": "Reconcile March invoices",
    "plan": ["fetch invoices", "match payments", "flag mismatches"],
    "completed": ["fetch invoices"],
    "facts": {"invoice_count": 142, "currency": "GBP"},
}

This makes runs resumable after a crash, lets you pause for approval and resume, and gives evaluators something concrete to assert against. See /courses/04-agentic-ai/lesson-04/ for persistence.

7. Least Privilege

Give the agent the minimum capability set for the task; pass a per-task allow-list, not the full registry. A summarization agent needs no write tools; a research agent needs no shell. Sandbox execution where possible.

Capability Grant only if the task requires it
File write Producing artifacts on disk
Shell / exec No safer narrow tool exists
Network egress Calling known, allow-listed hosts
Credentials Scoped, rotatable, least-permission tokens

Related References