Debugging & Error Resolution

45 min intermediate Lesson 5

Learning Outcomes

  • Share error messages and stack traces with Codex for accurate diagnosis
  • Use iterative fix-and-test cycles to resolve bugs systematically
  • Apply test-driven debugging to prevent regressions
  • Distinguish between errors Codex can fix directly and those requiring human judgment
  • Build a debugging workflow that combines Codex analysis with manual verification

Lesson Plan

Segment Duration Topic
Intro 3 min Debugging as a conversation — the Codex advantage
Explain 5 min How to share errors effectively
Demo 1 8 min Diagnosing a runtime error from a stack trace
Demo 2 8 min Iterative fixing — first attempt, test, refine
Demo 3 7 min Test-driven debugging workflow
Explain 5 min When Codex can't help — recognizing limits
Demo 4 6 min Complex multi-file bug investigation
Wrap-up 3 min Building your debugging toolkit

Before You Begin

Pre-work:

  • Complete Lesson 4 — Multi-Step Tasks
  • Have a project with at least one known bug (or introduce one intentionally)
  • Install a test runner for your language (pytest, jest, etc.)

Shopping List:

  • Codex CLI installed and configured
  • A project with a bug to fix (or use the example below)
  • A test runner installed (pytest, jest, mocha, etc.)
  • Familiarity with reading stack traces in your language

1 Sharing Errors Effectively

The quality of Codex's debugging help depends entirely on how you present the problem. A vague "it doesn't work" gets a vague answer. A precise error report gets a precise fix.

The ideal error report includes:

  1. What you were trying to do (context)
  2. The exact error message (copy-paste, don't paraphrase)
  3. The relevant file (so Codex can read it)
  4. Steps to reproduce (if not obvious)

Good error sharing:

> I'm running `python src/main.py` and getting this error:
>
> Traceback (most recent call last):
>   File "src/main.py", line 45, in process_order
>     total = calculate_total(items)
>   File "src/pricing.py", line 12, in calculate_total
>     price = item['price'] * item['quantity']
> TypeError: unsupported operand type(s) for *: 'str' and 'int'
>
> It happens when processing orders from the CSV import.

Bad error sharing:

> The calculate function is broken, can you fix it?

Sharing errors from the terminal:

You can run commands directly and Codex will see the output:

> Run the tests and show me any failures

Codex executes the test command, sees the full output (including errors), and can diagnose based on complete information.

# Let Codex run the failing test directly
codex "Run pytest tests/test_pricing.py -v and fix any failures"
# Let Codex run the failing test directly
codex "Run pytest tests/test_pricing.py -v and fix any failures"
TIP
Tip
Copy the complete error output — don't truncate it. The line numbers, file paths, and full traceback all give Codex essential context for accurate diagnosis.
NOTE
How It Works
When you paste an error message, Codex reads the referenced files (e.g., pricing.py line 12) to understand the surrounding code. This is why file paths in the error are so valuable.

2 Stack Trace Analysis

Codex excels at reading stack traces — it can follow the call chain, identify the root cause, and propose fixes at the right level.

Example: Python TypeError

> Here's the error from my order processing:
>
> Traceback (most recent call last):
>   File "src/main.py", line 45, in process_order
>     total = calculate_total(items)
>   File "src/pricing.py", line 12, in calculate_total
>     price = item['price'] * item['quantity']
> TypeError: unsupported operand type(s) for *: 'str' and 'int'
>
> Diagnose this and suggest a fix.

What Codex does:

  1. Reads src/pricing.py to see the calculate_total function
  2. Reads src/main.py to see where items comes from
  3. Identifies that item['price'] is a string (likely from CSV parsing)
  4. Proposes a fix — either type conversion at the source or in the function

Codex's analysis might look like:

The error occurs because item['price'] is a string, not a number.
This likely happens because the CSV import reads all values as strings.

I recommend fixing this at the data loading layer rather than in
calculate_total. Here's the fix for src/data_loader.py:
  def load_orders(filepath):
      with open(filepath) as f:
          reader = csv.DictReader(f)
-         return list(reader)
+         orders = []
+         for row in reader:
+             row['price'] = float(row['price'])
+             row['quantity'] = int(row['quantity'])
+             orders.append(row)
+         return orders

JavaScript example:

> Getting this in my Express app:
>
> TypeError: Cannot read properties of undefined (reading 'id')
>     at getUserProfile (/src/controllers/user.js:28:35)
>     at Layer.handle [as handle_request] (/node_modules/express/...)
>
> The endpoint is GET /api/users/:userId/profile

Codex will read the controller, identify that req.params.userId isn't being properly used to fetch the user (likely the database query returns undefined), and suggest adding a null check.

Multi-level stack traces:

For long stack traces, tell Codex which layer matters:

> Here's a long stack trace. The error originates in my code (not
> in node_modules). Focus on the application-level frames:
> [paste trace]
WARNING
Watch Out
Codex fixes the immediate error, but that might not be the root cause. If `item['price']` is a string, the real question is: why isn't it being parsed correctly at import time? Ask Codex to trace back to the source.

3 Iterative Fix-and-Test Cycles

Debugging rarely works on the first try. The iterative cycle — fix, test, observe, fix again — is where Codex shines because it remembers previous attempts.

The iterative debugging loop:

Attempt 1:
> Fix the TypeError in pricing.py by converting price to float

[approve the fix]

> Now run the tests

[Codex runs tests, sees a new error]

Attempt 2:
> The tests show a ValueError when price is "N/A". Handle that case too.

[approve the fix]

> Run the tests again

[all pass]

Practical example — fixing a failing test suite:

> Run pytest and show me all failures

Codex runs pytest and sees:

FAILED tests/test_api.py::test_create_user - AssertionError: 400 != 201
FAILED tests/test_api.py::test_list_users - KeyError: 'users'
FAILED tests/test_api.py::test_delete_user - 404 Not Found
> Fix the test_create_user failure first. Read the test to understand
> what it expects, then fix the implementation.

After that fix:

> Run pytest again. Did fixing create_user also fix list_users?

Continue until all tests pass.

When the fix introduces new bugs:

> The fix for the price parsing broke the export function.
> Run pytest tests/test_export.py to see what failed, then fix
> the export without reverting the price fix.

Codex can hold both requirements in mind and find a solution that satisfies both.

TIP
Tip
After each fix, run the full test suite — not just the failing test. Fixes often have ripple effects. Let Codex see the complete picture.

Keeping track of attempts:

> We've tried two approaches to fix the date parsing bug:
> 1. Using dateutil.parser (failed on timezone-aware dates)
> 2. Using datetime.strptime with explicit format (failed on
>    European date formats)
> Suggest a third approach that handles both cases.
NOTE
How It Works
Codex remembers previous failed attempts in the session context. It won't suggest the same fix twice unless you start a new session. Use this to your advantage by explicitly noting what didn't work.

4 Test-Driven Debugging

Test-driven debugging (TDD debugging) means writing a test that reproduces the bug BEFORE fixing it. This ensures your fix actually addresses the problem and prevents regressions.

The TDD debugging workflow:

Step 1: Write a failing test that reproduces the bug
Step 2: Confirm the test fails for the right reason
Step 3: Fix the code
Step 4: Confirm the test passes
Step 5: Check that no other tests broke

Example — reproducing a bug with a test:

> I found a bug: when a user's email contains a plus sign
> (like [email protected]), the registration fails silently.
> First, write a test in tests/test_registration.py that
> reproduces this bug.

Codex creates:

def test_registration_with_plus_in_email():
    """Regression test: emails with + should be accepted."""
    response = client.post("/register", json={
        "email": "[email protected]",
        "password": "SecurePass123!"
    })
    assert response.status_code == 201
    assert response.json()["email"] == "[email protected]"
> Run this test to confirm it fails

Output: FAILED - AssertionError: 400 != 201

> Good, the test reproduces the bug. Now fix the registration
> code to accept plus signs in emails.

After the fix:

> Run the full test suite to verify the fix and check for regressions

Why this approach is superior:

  1. You have proof the bug existed (the test)
  2. You have proof the fix works (the test passes)
  3. The bug can never come back undetected (regression test)
  4. You can show exactly what was wrong to reviewers
TIP
Tip
Ask Codex to write the test FIRST, before you even look at the fix. This keeps you honest — the test should fail before the fix, and pass after. If it passes before, the test doesn't actually reproduce the bug.

Edge case testing:

After fixing a bug, ask Codex to think about related edge cases:

> We fixed the plus-sign issue. What other special characters
> could cause similar problems in email validation? Write tests
> for those too.
WARNING
Watch Out
Make sure the test fails for the RIGHT reason. If it fails because of a setup issue (missing fixture, wrong import) rather than the actual bug, the test is not valid. Always check the failure message.

5 Multi-File Bug Investigation

Some bugs span multiple files — the symptom appears in one file, but the cause is in another. Codex's ability to read across your project makes it excellent at these investigations.

Investigation workflow:

> I'm getting incorrect totals in the invoice PDF. The displayed
> total doesn't match the sum of line items. Help me trace where
> the calculation goes wrong.
>
> The flow is: order.js → invoice_calculator.js → pdf_generator.js

Codex will read all three files and trace the data flow:

I've traced the issue. Here's what's happening:

1. order.js correctly calculates line_total = price × quantity
2. invoice_calculator.js applies a 10% tax, but it's using
   Math.round() which loses precision on large orders
3. pdf_generator.js sums the rounded values instead of
   rounding the final sum

The fix should be in invoice_calculator.js — delay rounding
until after summation.

Asking Codex to investigate:

> Something is corrupting the user session after login. The user
> object is correct at the login endpoint but missing fields by
> the time it reaches the dashboard endpoint. Trace the middleware
> chain and find where the session data gets modified.

Using Codex to add diagnostic logging:

> Add temporary debug logging to trace the request through the
> middleware chain. Log the user object at each stage so I can
> see where it changes. Mark these with "DEBUG-TRACE" so I can
> remove them later.

After finding the bug:

> Found it — the session serializer strips custom fields. Fix
> the serializer, then remove all "DEBUG-TRACE" log lines.

Cross-referencing with git history:

> This bug was introduced recently. Run git log --oneline -10
> and check which recent commits touched the session middleware.

Codex can correlate timing information with code changes to narrow down when a bug was introduced.

NOTE
How It Works
Codex investigates by reading files, running grep commands, and tracing function calls. In untrusted policy, you'll see these read operations happening. It builds a mental model of your code flow, similar to how a human debugger would.
TIP
Tip
For complex bugs, start by asking Codex to explain the data flow before asking for a fix. Understanding the architecture first leads to better fixes. Say 'Don't fix anything yet — just explain how the data flows from X to Y.'

Questions & Answers

Q: Can Codex debug runtime issues it can't reproduce (e.g., race conditions)?
Codex can analyze code for potential race conditions and suggest fixes based on patterns (adding locks, using atomic operations), but it cannot actually reproduce timing-dependent bugs. For these, describe the symptoms clearly and ask Codex to identify potential race conditions in the relevant code sections.
Q: How does Codex handle errors in languages it's less familiar with?
Codex (powered by OpenAI models) is strongest in Python, JavaScript/TypeScript, Java, C/C++, Go, and Rust. For less common languages, it can still analyze error messages and suggest fixes, but accuracy may be lower. Always verify suggestions more carefully for niche languages.
Q: Should I let Codex run my application in never policy to find bugs?
For running tests, on-request policy is reasonable (it can edit files but you approve commands). Full-auto is better reserved for CI-like environments. For interactive applications (web servers, GUIs), it's better to run the app yourself and paste errors into Codex, since Codex can't interact with running UIs.
Q: What if Codex's fix makes things worse?
This happens, especially with complex bugs. Revert the change (git checkout), provide the new error information to Codex, and explicitly state what the previous fix broke. Codex learns from this context within the session and will try a different approach. If it keeps failing, consider breaking the problem down further or researching the specific error manually.

Key Takeaways

  1. Quality in, quality out: The better you describe the error, the better the fix — always include the full error output
  2. Iterate systematically: Fix one thing, test, observe — don't try to fix everything at once
  3. Test-driven debugging: Write a failing test first, then fix — this prevents regressions
  4. Multi-file investigation: Ask Codex to trace data flows across files before jumping to fixes
  5. Know the limits: Codex can't reproduce race conditions or interact with running UIs — use it for analysis and code fixes
  6. Git protects you: Always be able to revert if a fix attempt makes things worse

Next Steps: In Lesson 6 — Project Scaffolding, you'll learn how to use Codex to generate complete project structures from scratch.