Task Decomposition at Scale

45 min intermediate Lesson 4

Learning Outcomes

  • Classify a complex goal into functional, data, temporal, or hybrid decomposition strategies
  • Model sub-task dependencies as a DAG and compute parallel versus sequential stages
  • Build a supervisor that dynamically decomposes work based on complexity estimates
  • Pass context between sub-tasks idempotently so retries and partial failures stay safe
  • Reassemble distributed sub-results into one coherent output with conflict handling

Lesson Plan

Segment Duration Topic
Intro 3 min Why decomposition is the load-bearing wall of orchestration
Concepts 6 min The four decomposition strategies
Build 8 min A typed plan: tasks, dependencies, the DAG
Build 8 min Dynamic decomposition by complexity
Build 8 min Scheduling: topological waves and parallelism
Build 7 min Result reassembly and conflict resolution
Wrap-up 5 min Depth trade-offs, takeaways, preview

Before You Begin

Pre-work:

Shopping List:

  • Python 3.11+ with pip install anthropic
  • An ANTHROPIC_API_KEY exported in your environment
  • A graph mental model: DAGs, topological ordering, fan-out/fan-in

1 The Four Decomposition Strategies

A single agent fails complex goals for the reasons in Lesson 1: finite context, no parallelism, one prompt trying to be five specialists at once. Decomposition cuts a goal into units small enough that each fits one agent's context and skill set. Four axes to cut along:

Strategy Cut along… Example Shape
Functional Skill / role extractor → validator → summarizer Pipeline / supervisor
Data Partition of input 10,000 invoices into shards Fan-out / fan-in
Temporal Phase over time plan → build → test → review Sequential gated stages
Hybrid Two+ axes at once per-region × per-doc-type A DAG, not a line

Functional is splitting a monolith into microservices by responsibility; data is sharding; temporal is a staged pipeline where each phase gates the next. Real systems are almost always hybrid — fan out searches by source (data), route each through specialist analyzers (functional), merge in a synthesis phase (temporal).

NOTE
Key Insight
The strategy sets your failure surface. Data decomposition gives N independent failures you retry in isolation; temporal decomposition gives a serial chain where any phase blocks the whole run. Choose the cut that makes failures cheap to recover from.

2 A Typed Plan: Tasks and the Execution DAG

Decomposition output should be a plan — a serializable structure, not free text. A plan is a set of tasks plus dependency edges; the edges form a directed acyclic graph (DAG), and the absence of cycles is what lets you schedule it.

from dataclasses import dataclass, field

@dataclass(frozen=True)
class Task:
    id: str                  # stable, content-derived
    agent: str               # which specialist handles this
    instruction: str         # self-contained unit of work
    depends_on: tuple = ()    # ids that must finish first
    est_tokens: int = 4000   # complexity estimate, drives budgeting

@dataclass
class Plan:
    goal: str
    tasks: dict = field(default_factory=dict)   # id -> Task

    def validate(self) -> None:
        for t in self.tasks.values():
            for dep in t.depends_on:
                if dep not in self.tasks:
                    raise ValueError(f"{t.id} needs unknown {dep}")
        assert_acyclic(self.tasks)   # DFS that raises on a back-edge

The acyclicity check is non-negotiable: a cycle is a deadlock a workflow engine runs until your timeout budget is exhausted. Implement assert_acyclic as a DFS coloring nodes white/gray/black that raises the moment it revisits a gray (in-progress) node. Run validate() before any agent fires, and treat the Plan as the contract the supervisor produces, a guardrail inspects, and the scheduler consumes.

WARNING
Watch Out
Make task ids deterministic — derive them from a hash of agent plus instruction, not a random UUID or timestamp. Stable ids make result caching and idempotent retries possible. A new id every run means you cache nothing and replay nothing.

3 Dynamic Decomposition by Complexity Estimate

At scale you want the supervisor agent to decompose dynamically. Have it emit the plan as structured JSON via tool use, never as prose you parse with a regex. Define a Task JSON Schema (the same fields as Step 2's dataclass, with agent as an enum of your real specialists), wrap it in a tasks array as one tool, and force the model to call it:

from anthropic import Anthropic
client = Anthropic()

def decompose(goal: str, plan_tool: dict) -> list[dict]:
    msg = client.messages.create(
        model="claude-sonnet-4-6", max_tokens=2048,
        tools=[plan_tool],
        tool_choice={"type": "tool", "name": "emit_plan"},
        system=("Decompose the goal into the smallest set of independent "
                "sub-tasks. Prefer data-parallel fan-out. Add a depends_on "
                "edge only when a task truly needs another's output."),
        messages=[{"role": "user", "content": goal}],
    )
    block = next(b for b in msg.content if b.type == "tool_use")
    return block.input["tasks"]

Forcing tool_choice guarantees valid JSON in your schema; the enum on agent stops the supervisor inventing a specialist you never wired up. Add two guardrails: cap plan size (reject decompositions over ~50 tasks — a runaway planner is a runaway bill) and bound recursion depth to 2–3 levels, since each level of further decomposition multiplies cost.

TIP
Tip
Keep the planner's system prompt focused only on planning, at low temperature. Mixing planning and execution produces a supervisor that does the work itself instead of delegating — a manager who keeps writing the code.
WARNING
Watch Out
The model's est_tokens is a guess. Treat it as a prior for budgeting and routing, then reconcile against real usage. Lesson 9 turns these into hard token budgets per agent.

4 Scheduling: Topological Waves and Parallelism

Given a validated DAG, the scheduler repeatedly asks: what can run right now? A task is ready when all dependencies have completed. Group ready tasks into waves dispatched concurrently — a level-order (Kahn's algorithm) topological sort, fan out each wave then fan in before the next.

import asyncio

def waves(plan: Plan) -> list[list[str]]:
    remaining, done, out = dict(plan.tasks), set(), []
    while remaining:
        ready = [tid for tid, t in remaining.items()
                 if set(t.depends_on) <= done]
        if not ready:
            raise RuntimeError("no progress — cycle or missing dep")
        out.append(ready)
        for tid in ready:
            del remaining[tid]
        done.update(ready)
    return out

async def run_plan(plan: Plan) -> dict:
    ctx = {}
    for wave in waves(plan):
        async def run_one(tid):
            t = plan.tasks[tid]
            inputs = {d: ctx[d] for d in t.depends_on}  # only what it needs
            return tid, await execute_agent(t, inputs)
        pairs = await asyncio.gather(*(run_one(t) for t in wave),
                                     return_exceptions=True)
        ctx.update({k: v for k, v in pairs if not isinstance(v, Exception)})
    return ctx

Each agent receives only its declared dependencies' outputs, not the whole accumulated context — a fan-out task over invoice #4,712 never sees invoice #11's analysis, which keeps every prompt small and free of cross-contamination. Because ids are stable, wrap execute_agent in a cache lookup so a retried wave re-runs only the tasks that failed.

NOTE
Key Insight
Wave width is your parallelism, and parallelism costs money and rate-limit budget at once. A 200-task fan-out as one wave hits your concurrency limit and your wallet simultaneously. Bound wave width with a semaphore and drain it in chunks.
WARNING
Watch Out
Plain asyncio.gather raises on the first exception and cancels the rest. Pass return_exceptions=True so one failed shard does not vaporize the nine that succeeded — full recovery is Lesson 7, but design for it now.

5 Result Reassembly and Conflict Resolution

Decomposition is half the job; results must come back together, and the reassembly strategy mirrors how you cut the work.

Decomposition Reassembly pattern Watch for
Functional (pipeline) Last stage's output is the answer Lost intermediate evidence
Data (fan-out) Concatenate / aggregate / reduce Duplicates, conflicts
Temporal (phases) Final phase plus audit trail Stale data from earlier phases

Data fan-out is where reassembly gets interesting, because independent agents can disagree — two shards might extract a different "total" for the same record. Run a deterministic merge that unions records and resolves what it can (dedupe, sum, take-latest), then escalate genuine semantic conflicts to a reconciliation agent with just the conflicting values and their provenance:

> You are a reconciliation agent. Two extractors disagree on a field.
> Field: invoice_total
> Value A: 4820.00  (source: shard_03, page 2)
> Value B: 4280.00  (source: shard_07, page 2)
> Return the correct value and a one-line justification. If you cannot
> determine it, return needs_human: true.

The needs_human: true exit is deliberate. Reassembly is where eventual consistency bites — agents ran at different times against possibly different context, so forcing a confident merge on ambiguous data manufactures plausible-looking wrong answers. A clean escalation path is cheaper than a confidently incorrect total.

TIP
Tip
Preserve provenance through reassembly — every merged value should carry which task or shard produced it. It turns 'where did this number come from?' into a one-line lookup, and underpins the observability you add in Lesson 8.

6 Choosing Decomposition Depth — The Trade-off

More decomposition is not strictly better. Each sub-task carries fixed overhead: a fresh prompt, re-sent context, a round-trip, a slice of concurrency budget. Too coarse and an agent chokes on a task too big for its context; too fine and coordination overhead dominates the real work. Encode the stop rule explicitly — a should_decompose(task, depth) predicate that returns False once depth hits a max_depth of 3 or task.est_tokens falls under a single-agent threshold (say 6,000), and only otherwise splits further.

Signal Lean coarser Lean finer
Task fits one context window yes —
Sub-tasks embarrassingly parallel — yes (fan out)
Strong inter-task dependencies yes —
Need independent retry of pieces — yes
Different specialist tools per piece — yes

Every dependency edge is a synchronization point and a place a failure can stall the DAG; every wider wave trades latency for throughput against your rate limit. Depth is a tuning knob between coordination cost and single-agent overload — no free lunch, only the trade-off that fits this workload's SLOs.

NOTE
Key Insight
A good plan is the minimum decomposition that keeps every task inside one agent's competence and context. Start coarse, measure where tasks overflow or fail, and split only those. Premature fan-out is the multi-agent version of premature optimization.

Questions & Answers

Q: How do I stop a dynamic planner from producing a wildly different plan each run and breaking reproducibility?
Force structured output with a tool schema (Step 3), set low or zero temperature so the same goal yields the same plan, and derive task ids deterministically from content so identical work maps to identical ids. For audit-grade reproducibility, persist the emitted plan and replay from the stored plan rather than re-planning.
Q: My data fan-out has 5,000 shards. Do I really dispatch one agent call each?
Not naively. Batch small records into one call until you approach a sensible context budget, and bound concurrency with a semaphore so the wave drains in controlled chunks instead of slamming your rate limit. Also ask whether an LLM is even right per shard — if the per-shard work is deterministic, do it in code and reserve agents for genuinely ambiguous shards. Cost optimization for this is the subject of Lesson 9.
Q: One task in a wave fails but its dependents are already scheduled — what happens?
They are not — the scheduler marks a task ready only when all dependencies are in the done set, and a failed task never enters it. So failures quarantine their downstream subtree, which stalls rather than corrupts. What to do with the stalled subtree (retry, fallback agent, degraded path, escalate) is Lesson 7. For now, ensure gather uses return_exceptions=True so the rest of the wave completes and you capture which task failed.
Q: Passing only declared dependencies sounds clean, but what about context every task needs — the goal, global config?
Separate two kinds. Global, read-only context (goal, account settings, schema) is injected into every task's system prompt as immutable preamble. Task-specific context (another task's output) flows only through declared depends_on edges. Anything an agent can change or that is branch-specific travels through edges; anything global and read-only is ambient. Mixing them causes the blackboard contention covered in Lesson 3.
Q: How deep should recursive decomposition go before it is just turtles all the way down?
In practice 2 to 3 levels covers almost everything: a top supervisor, optional mid-level supervisors per branch, and leaf workers. Each level multiplies coordination cost and adds a layer where failures stall, so treat depth as a hard cap (see should_decompose). Needing more than three levels usually signals the goal should split into separate workflows that hand off, not one mega-DAG.

Key Takeaways

  1. Four cuts, usually combined — functional (by skill), data (by partition), temporal (by phase), and hybrid. The cut sets your parallelism and your failure surface.
  2. The plan is a typed DAG, not prose — emit a validated, acyclic, inspectable structure with stable task ids so it can be approved, cached, and replayed.
  3. Decompose dynamically with forced tool use — size the plan via a JSON tool schema, then cap plan size and recursion depth to bound cost.
  4. Schedule in topological waves — a task is ready when its dependencies finish; dispatch each wave concurrently, pass each agent only its declared inputs, and bound wave width.
  5. Reassembly mirrors decomposition — pipelines take the last stage, fan-outs reduce with explicit conflict handling, and semantic conflicts escalate to a reconciliation agent or a human.
  6. Depth is a trade-off, not a virtue — the right plan is the minimum decomposition that keeps every task inside one agent's context and competence.

Next Steps: Lesson 5: The Anthropic Agent SDK