Multi-Step Tasks

45 min intermediate Lesson 4

Learning Outcomes

  • Structure complex requests that require multiple steps to complete
  • Use task decomposition to break large goals into manageable instructions
  • Chain file operations, shell commands, and edits in a logical sequence
  • Implement rollback strategies when multi-step tasks go wrong
  • Maintain context across a multi-step session effectively

Lesson Plan

Segment Duration Topic
Intro 3 min Why multi-step matters — real tasks are never one command
Explain 5 min Task decomposition strategies
Demo 1 9 min Chaining a 4-step feature addition
Demo 2 8 min Complex refactoring with dependencies
Explain 5 min Context management across steps
Demo 3 7 min Rollback when step 3 of 5 fails
Demo 4 5 min Using git checkpoints within a session
Wrap-up 3 min Patterns for reliable multi-step tasks

Before You Begin

Pre-work:

  • Complete Lesson 3 — File Operations
  • Set up a small project with at least 5 files (e.g., a simple web app or CLI tool)
  • Ensure git is initialized with a clean commit

Shopping List:

  • Codex CLI installed and configured
  • A project with interconnected files (imports between modules)
  • Git initialized with at least one commit
  • Comfort with the approval flow from Lessons 1-3

1 Understanding Multi-Step Tasks

Real coding tasks almost never involve just one file or one change. Adding a feature might require creating a module, updating imports, modifying a config, and adding tests. Codex CLI handles these naturally through conversation.

Single-step vs multi-step:

Single-Step Multi-Step
"Add a docstring to this function" "Add a user registration endpoint with validation, database model, and tests"
"Fix the typo in line 12" "Refactor the auth module to use JWT instead of session tokens"
"Create a .gitignore file" "Set up a complete CI/CD pipeline with linting, testing, and deployment"

How Codex handles multi-step tasks:

When you give a complex instruction, Codex:

  1. Analyzes the full scope of what's needed
  2. Plans a sequence of operations
  3. Executes them one at a time, seeking approval for each
  4. Maintains context — each step knows what came before

Two approaches to multi-step work:

Approach A — Single detailed prompt:

> Add a new /api/products endpoint. This needs:
> 1. A Product model in src/models/product.js
> 2. A route handler in src/routes/products.js
> 3. Register the route in src/app.js
> 4. Add validation middleware
> 5. Write basic tests

Approach B — Iterative conversation:

> Create a Product model in src/models/product.js with name, price, and description fields
[approve]
> Now create a route handler for /api/products with GET and POST
[approve]
> Register this route in app.js
[approve]
> Add input validation for the POST endpoint
[approve]
> Write tests for the new endpoints
[approve]
TIP
Tip
Approach A is faster when you know exactly what you want. Approach B gives you more control and lets you course-correct between steps. Use B when you're exploring or unsure about implementation details.

2 Task Decomposition Strategies

Breaking a large task into logical steps is a skill that makes multi-step Codex sessions more reliable. Here are proven decomposition strategies.

Strategy 1: Dependency order (bottom-up)

Start with the parts that have no dependencies, then build up:

> Step 1: Create the database schema for products
> Step 2: Create the data access layer (uses schema)
> Step 3: Create the business logic layer (uses data access)
> Step 4: Create the API endpoint (uses business logic)
> Step 5: Create tests (uses the endpoint)

Strategy 2: File-by-file

When changes span many files, go one file at a time:

> Let's refactor the error handling. Start with src/api/users.js —
> replace all try-catch blocks with the centralized error handler.

Then move to the next file in a follow-up prompt.

Strategy 3: Interface first, implementation second

Define the contract before filling in details:

> Create src/services/payment.js with function stubs (name,
> parameters, return types, JSDoc) but no implementation yet.

Then:

> Now implement the processPayment function using Stripe's API.

Strategy 4: Test first (TDD)

Write the test, then implement:

> Write a test in tests/test_calculator.py that tests an add()
> function from src/calculator.py. The function doesn't exist yet.

Then:

> Now create src/calculator.py with an add() function that makes
> the test pass.
NOTE
How It Works
Codex remembers all previous steps in the conversation. When you say 'now implement the function', it knows which function you're referring to because it created the stub earlier in the session.

Decomposition checklist:

Before starting a complex task, ask yourself:

  • What files need to be created?
  • What files need to be modified?
  • What's the dependency order?
  • Which step, if wrong, would cascade failures?
  • Where should I add git checkpoints?

3 Chaining Operations in Practice

Let's walk through a realistic multi-step task: adding a complete feature to an existing Express.js application.

The goal: Add a "forgot password" flow with email sending.

Step 1 — Create the email utility:

> Create src/utils/email.js with a sendEmail function that uses
> nodemailer. Include a sendPasswordReset helper that takes an
> email address and a reset token.

After approval, the file is created.

Step 2 — Add the route:

> Create src/routes/passwordReset.js with two endpoints:
> POST /forgot-password (accepts email, generates token, sends email)
> POST /reset-password (accepts token and new password, updates user)

Codex reads the email utility it just created and uses it in the route.

Step 3 — Update the user model:

> Add resetToken and resetTokenExpiry fields to the User model
> in src/models/user.js

Step 4 — Register the route:

> Register the password reset routes in src/app.js, alongside
> the existing auth routes

Step 5 — Add tests:

> Write integration tests for both password reset endpoints in
> tests/passwordReset.test.js. Mock the email sending.

What happens during chaining:

At each step, Codex:

  • Reads relevant existing files to understand the codebase
  • Generates changes consistent with what already exists
  • References code created in earlier steps
  • Maintains naming conventions and patterns from your project
WARNING
Watch Out
If you reject a step and ask for something different, Codex adapts its context. But if you approve step 2 then realize step 1 was wrong, you'll need to go back and fix step 1 — Codex won't automatically reconcile earlier mistakes.

Verifying between steps:

You can ask Codex to verify work between steps:

> Before we continue, check that the imports in passwordReset.js
> are consistent with the email utility we created

4 Rollback Strategies

When multi-step tasks go wrong — a step produces incorrect code, or you realize the approach is flawed — you need rollback strategies.

Strategy 1: Git checkpoints

Create commits between major steps:

> Run: git add -A && git commit -m "feat: add email utility"

Then if step 3 goes wrong:

# Revert to the last checkpoint
git reset --hard HEAD~1

# Or revert just the files from the bad step
git checkout HEAD -- src/models/user.js
# Revert to the last checkpoint
git reset --hard HEAD~1

# Or revert just the files from the bad step
git checkout HEAD -- src/models/user.js

Strategy 2: Ask Codex to undo

Codex can revert its own changes:

> The last change to user.js was wrong. Revert it back to how it
> was before (the version without resetToken fields).

This works for simple reversals, but for complex multi-file changes, git is more reliable.

Strategy 3: Selective rejection

In untrusted policy, you can reject individual steps:

  1. Codex proposes changing file A — you approve
  2. Codex proposes changing file B — you reject
  3. Codex proposes an alternative for file B — you approve

Strategy 4: Fresh start with context

If things get too tangled:

> Let's start over on the password reset feature. First, revert
> all changes we've made (git reset --hard HEAD~3), then let me
> give you a revised plan.
TIP
Tip
The best rollback strategy is prevention. Commit after every successful step in a complex task. The command `git add -A && git commit -m 'checkpoint: step N done'` takes seconds and can save hours.

Recovery workflow after a failed step:

> The test file you just created has import errors. Let's fix it:
> 1. Check what imports are available from our modules
> 2. Update the imports in the test file
> 3. Verify the tests can at least be parsed without errors
WARNING
Watch Out
If you're in on-request policy and a step goes wrong, the files are already changed. Always have git checkpoints when using auto-edit for multi-step tasks.

5 Context Management Across Steps

Codex maintains conversation context throughout a session, but there are limits and best practices for keeping context effective.

What Codex remembers:

  • All messages you've sent in the session
  • All files it has read or created
  • The approval/rejection decisions you made
  • Shell command outputs

What can degrade context:

  • Very long sessions (context window fills up)
  • Switching between unrelated tasks
  • Large file contents consuming context space

Best practices for context management:

1. Keep related tasks in one session:

# Good — related steps stay together
Session 1: Build the password reset feature (steps 1-5)
Session 2: Build the user profile feature (steps 1-4)

# Bad — mixing unrelated work
Session 1: Password reset step 1, then profile step 1, then reset step 2...

2. Summarize before continuing:

If you're deep in a session, remind Codex of the state:

> So far we've created the email utility and the route handler.
> Next, let's update the user model to include the reset token.

3. Reference files explicitly:

> Looking at the sendPasswordReset function in src/utils/email.js,
> update the route handler to call it correctly.

4. Start fresh when context degrades:

If Codex starts producing inconsistent results:

> /clear

Then start a new conversation with a summary:

> I'm building a password reset feature. I already have:
> - src/utils/email.js (sendPasswordReset function)
> - src/routes/passwordReset.js (two endpoints)
> Now I need to add the resetToken field to the User model.

5. Use git log as external memory:

> Look at the last 3 git commits to understand what we've built
> so far, then continue with the next step.
NOTE
How It Works
Codex's context window has a finite size. For very long sessions (20+ exchanges), earlier messages may be summarized or truncated. Git commits serve as persistent memory that survives across sessions.
TIP
Tip
If a multi-step task spans more than 15-20 exchanges, consider starting a new session and providing context through existing files and git history rather than relying on conversation memory alone.

Questions & Answers

Q: What's the maximum number of steps Codex can handle in one session?
There's no hard limit on steps, but context quality degrades in very long sessions (30+ exchanges). The model's context window fills up and earlier details may be lost. For large tasks, break them into multiple sessions of 10-15 exchanges each, using git commits as checkpoints between sessions.
Q: Can I save a multi-step plan and resume it later?
Codex CLI doesn't have a built-in "save session" feature. However, you can write your plan to a file (e.g., PLAN.md) during the session, commit it, and reference it in a new session later. Ask Codex to "read PLAN.md and continue from step 4" in your next session.
Q: How does Codex handle dependencies between steps if one fails?
If a step fails (command error, or you reject it), Codex knows about the failure and adjusts its approach. It might suggest an alternative implementation or ask clarifying questions. It won't blindly proceed with steps that depend on a failed step — it'll flag the dependency issue.
Q: Should I use never policy for multi-step tasks?
Generally no, especially while learning. Multi-step tasks have more opportunities for errors to compound. Use untrusted policy so you can catch issues early. Once you've run a similar multi-step task successfully in untrusted policy and understand the pattern, you might use auto-edit for subsequent runs of the same type of task.

Key Takeaways

  1. Two approaches: Use detailed single prompts when you're sure of the plan; use iterative conversation when exploring
  2. Decompose by dependency: Start with independent pieces and build up to code that depends on them
  3. Git checkpoints are essential: Commit after each successful step so you can roll back precisely
  4. Context degrades over time: Keep sessions focused on related tasks, summarize state periodically
  5. Reject and retry: Don't approve steps you're unsure about — reject and rephrase instead
  6. Verify between steps: Ask Codex to check consistency before moving to the next step

Next Steps: In Lesson 5 — Debugging & Error Resolution, you'll learn how to use Codex CLI to diagnose and fix bugs efficiently.