Multi-Step Tasks
Learning Outcomes
- Structure complex requests that require multiple steps to complete
- Use task decomposition to break large goals into manageable instructions
- Chain file operations, shell commands, and edits in a logical sequence
- Implement rollback strategies when multi-step tasks go wrong
- Maintain context across a multi-step session effectively
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why multi-step matters — real tasks are never one command |
| Explain | 5 min | Task decomposition strategies |
| Demo 1 | 9 min | Chaining a 4-step feature addition |
| Demo 2 | 8 min | Complex refactoring with dependencies |
| Explain | 5 min | Context management across steps |
| Demo 3 | 7 min | Rollback when step 3 of 5 fails |
| Demo 4 | 5 min | Using git checkpoints within a session |
| Wrap-up | 3 min | Patterns for reliable multi-step tasks |
Before You Begin
Pre-work:
- Complete Lesson 3 — File Operations
- Set up a small project with at least 5 files (e.g., a simple web app or CLI tool)
- Ensure git is initialized with a clean commit
Shopping List:
- Codex CLI installed and configured
- A project with interconnected files (imports between modules)
- Git initialized with at least one commit
- Comfort with the approval flow from Lessons 1-3
Real coding tasks almost never involve just one file or one change. Adding a feature might require creating a module, updating imports, modifying a config, and adding tests. Codex CLI handles these naturally through conversation.
Single-step vs multi-step:
| Single-Step | Multi-Step |
|---|---|
| "Add a docstring to this function" | "Add a user registration endpoint with validation, database model, and tests" |
| "Fix the typo in line 12" | "Refactor the auth module to use JWT instead of session tokens" |
| "Create a .gitignore file" | "Set up a complete CI/CD pipeline with linting, testing, and deployment" |
How Codex handles multi-step tasks:
When you give a complex instruction, Codex:
- Analyzes the full scope of what's needed
- Plans a sequence of operations
- Executes them one at a time, seeking approval for each
- Maintains context — each step knows what came before
Two approaches to multi-step work:
Approach A — Single detailed prompt:
> Add a new /api/products endpoint. This needs:
> 1. A Product model in src/models/product.js
> 2. A route handler in src/routes/products.js
> 3. Register the route in src/app.js
> 4. Add validation middleware
> 5. Write basic tests
Approach B — Iterative conversation:
> Create a Product model in src/models/product.js with name, price, and description fields
[approve]
> Now create a route handler for /api/products with GET and POST
[approve]
> Register this route in app.js
[approve]
> Add input validation for the POST endpoint
[approve]
> Write tests for the new endpoints
[approve]
Breaking a large task into logical steps is a skill that makes multi-step Codex sessions more reliable. Here are proven decomposition strategies.
Strategy 1: Dependency order (bottom-up)
Start with the parts that have no dependencies, then build up:
> Step 1: Create the database schema for products
> Step 2: Create the data access layer (uses schema)
> Step 3: Create the business logic layer (uses data access)
> Step 4: Create the API endpoint (uses business logic)
> Step 5: Create tests (uses the endpoint)
Strategy 2: File-by-file
When changes span many files, go one file at a time:
> Let's refactor the error handling. Start with src/api/users.js —
> replace all try-catch blocks with the centralized error handler.
Then move to the next file in a follow-up prompt.
Strategy 3: Interface first, implementation second
Define the contract before filling in details:
> Create src/services/payment.js with function stubs (name,
> parameters, return types, JSDoc) but no implementation yet.
Then:
> Now implement the processPayment function using Stripe's API.
Strategy 4: Test first (TDD)
Write the test, then implement:
> Write a test in tests/test_calculator.py that tests an add()
> function from src/calculator.py. The function doesn't exist yet.
Then:
> Now create src/calculator.py with an add() function that makes
> the test pass.
Decomposition checklist:
Before starting a complex task, ask yourself:
- What files need to be created?
- What files need to be modified?
- What's the dependency order?
- Which step, if wrong, would cascade failures?
- Where should I add git checkpoints?
Let's walk through a realistic multi-step task: adding a complete feature to an existing Express.js application.
The goal: Add a "forgot password" flow with email sending.
Step 1 — Create the email utility:
> Create src/utils/email.js with a sendEmail function that uses
> nodemailer. Include a sendPasswordReset helper that takes an
> email address and a reset token.
After approval, the file is created.
Step 2 — Add the route:
> Create src/routes/passwordReset.js with two endpoints:
> POST /forgot-password (accepts email, generates token, sends email)
> POST /reset-password (accepts token and new password, updates user)
Codex reads the email utility it just created and uses it in the route.
Step 3 — Update the user model:
> Add resetToken and resetTokenExpiry fields to the User model
> in src/models/user.js
Step 4 — Register the route:
> Register the password reset routes in src/app.js, alongside
> the existing auth routes
Step 5 — Add tests:
> Write integration tests for both password reset endpoints in
> tests/passwordReset.test.js. Mock the email sending.
What happens during chaining:
At each step, Codex:
- Reads relevant existing files to understand the codebase
- Generates changes consistent with what already exists
- References code created in earlier steps
- Maintains naming conventions and patterns from your project
Verifying between steps:
You can ask Codex to verify work between steps:
> Before we continue, check that the imports in passwordReset.js
> are consistent with the email utility we created
When multi-step tasks go wrong — a step produces incorrect code, or you realize the approach is flawed — you need rollback strategies.
Strategy 1: Git checkpoints
Create commits between major steps:
> Run: git add -A && git commit -m "feat: add email utility"
Then if step 3 goes wrong:
# Revert to the last checkpoint
git reset --hard HEAD~1
# Or revert just the files from the bad step
git checkout HEAD -- src/models/user.js
# Revert to the last checkpoint
git reset --hard HEAD~1
# Or revert just the files from the bad step
git checkout HEAD -- src/models/user.js
Strategy 2: Ask Codex to undo
Codex can revert its own changes:
> The last change to user.js was wrong. Revert it back to how it
> was before (the version without resetToken fields).
This works for simple reversals, but for complex multi-file changes, git is more reliable.
Strategy 3: Selective rejection
In untrusted policy, you can reject individual steps:
- Codex proposes changing file A — you approve
- Codex proposes changing file B — you reject
- Codex proposes an alternative for file B — you approve
Strategy 4: Fresh start with context
If things get too tangled:
> Let's start over on the password reset feature. First, revert
> all changes we've made (git reset --hard HEAD~3), then let me
> give you a revised plan.
Recovery workflow after a failed step:
> The test file you just created has import errors. Let's fix it:
> 1. Check what imports are available from our modules
> 2. Update the imports in the test file
> 3. Verify the tests can at least be parsed without errors
Codex maintains conversation context throughout a session, but there are limits and best practices for keeping context effective.
What Codex remembers:
- All messages you've sent in the session
- All files it has read or created
- The approval/rejection decisions you made
- Shell command outputs
What can degrade context:
- Very long sessions (context window fills up)
- Switching between unrelated tasks
- Large file contents consuming context space
Best practices for context management:
1. Keep related tasks in one session:
# Good — related steps stay together
Session 1: Build the password reset feature (steps 1-5)
Session 2: Build the user profile feature (steps 1-4)
# Bad — mixing unrelated work
Session 1: Password reset step 1, then profile step 1, then reset step 2...
2. Summarize before continuing:
If you're deep in a session, remind Codex of the state:
> So far we've created the email utility and the route handler.
> Next, let's update the user model to include the reset token.
3. Reference files explicitly:
> Looking at the sendPasswordReset function in src/utils/email.js,
> update the route handler to call it correctly.
4. Start fresh when context degrades:
If Codex starts producing inconsistent results:
> /clear
Then start a new conversation with a summary:
> I'm building a password reset feature. I already have:
> - src/utils/email.js (sendPasswordReset function)
> - src/routes/passwordReset.js (two endpoints)
> Now I need to add the resetToken field to the User model.
5. Use git log as external memory:
> Look at the last 3 git commits to understand what we've built
> so far, then continue with the next step.
Questions & Answers
Key Takeaways
- Two approaches: Use detailed single prompts when you're sure of the plan; use iterative conversation when exploring
- Decompose by dependency: Start with independent pieces and build up to code that depends on them
- Git checkpoints are essential: Commit after each successful step so you can roll back precisely
- Context degrades over time: Keep sessions focused on related tasks, summarize state periodically
- Reject and retry: Don't approve steps you're unsure about — reject and rephrase instead
- Verify between steps: Ask Codex to check consistency before moving to the next step
Next Steps: In Lesson 5 — Debugging & Error Resolution, you'll learn how to use Codex CLI to diagnose and fix bugs efficiently.