Testing with Codex

45 min intermediate Lesson 7

Learning Outcomes

  • Generate comprehensive unit tests for existing code with proper assertions
  • Analyze code coverage reports and identify gaps to fill
  • Implement a TDD workflow using Codex to write tests before implementation
  • Fix failing tests by diagnosing whether the test or implementation is wrong
  • Generate edge case tests, integration tests, and mock-based tests
WARNING
Important
AI-generated tests can pass while code is still buggy — if the test makes the same wrong assumption as the implementation, both will agree on the wrong answer. Always review test assertions critically: do they verify the RIGHT behavior, not just ANY behavior?

Lesson Plan

Segment Duration Topic
Intro 3 min Why AI-generated tests are transformative
Demo 1 8 min Generating unit tests for an existing module
Demo 2 7 min Coverage analysis — finding and filling gaps
Explain 5 min TDD workflow with Codex
Demo 3 8 min Red-green-refactor with Codex assistance
Demo 4 7 min Fixing failing tests — test vs implementation bugs
Explain 4 min Mocking, fixtures, and test isolation
Wrap-up 3 min Building a test-first habit

Before You Begin

Pre-work:

  • Complete Lesson 6 — Project Scaffolding
  • Have a project with testable code (functions that take inputs and return outputs)
  • Install a test runner and coverage tool for your language

Shopping List:

  • Codex CLI installed and configured
  • A project with at least 3 untested modules
  • Test runner installed: pytest (Python), jest (JS/TS), or equivalent
  • Coverage tool: pytest-cov, jest --coverage, or equivalent
  • Familiarity with basic testing concepts (assertions, fixtures)

1 Generating Unit Tests

Codex can generate comprehensive test suites by reading your source code and creating tests for every function, method, and edge case.

Basic test generation:

> Read src/utils/validators.py and generate a complete test file
> at tests/test_validators.py. Test every function with at least
> 3 cases each: valid input, invalid input, and edge case.

Codex reads the source file, identifies all public functions, and generates:

import pytest
from src.utils.validators import validate_email, validate_phone, validate_url


class TestValidateEmail:
    def test_valid_standard_email(self):
        assert validate_email("[email protected]") is True

    def test_valid_email_with_plus(self):
        assert validate_email("[email protected]") is True

    def test_invalid_missing_at(self):
        assert validate_email("userexample.com") is False

    def test_invalid_empty_string(self):
        assert validate_email("") is False

    def test_edge_case_unicode(self):
        assert validate_email("user@examplé.com") is False


class TestValidatePhone:
    def test_valid_us_number(self):
        assert validate_phone("+11234567890") is True

    def test_valid_international(self):
        assert validate_phone("+447911123456") is True

    def test_invalid_too_short(self):
        assert validate_phone("123") is False

    def test_invalid_letters(self):
        assert validate_phone("abc-def-ghij") is False

    def test_edge_case_empty(self):
        assert validate_phone("") is False

Specifying test style:

> Generate tests for src/services/order.js using Jest.
> Use describe/it blocks, follow AAA pattern (Arrange, Act, Assert),
> and add descriptive test names that read like documentation.

Testing classes and methods:

> Read src/models/ShoppingCart.ts and generate tests for all public
> methods. Include tests for:
> - Empty cart behavior
> - Adding single and multiple items
> - Removing items that exist vs don't exist
> - Calculating totals with and without discounts
> - Quantity limits and validation
# Generate tests and immediately run them
codex "Generate tests for src/utils/validators.py, save to tests/test_validators.py, then run pytest tests/test_validators.py -v"
# Generate tests and immediately run them
codex "Generate tests for src/utils/validators.py, save to tests/test_validators.py, then run pytest tests/test_validators.py -v"
TIP
Tip
Ask Codex to run the tests immediately after generating them. This catches issues early — incorrect imports, wrong function signatures, or implementation bugs.
WARNING
Watch Out
Generated tests may pass even when the code is buggy if the test makes the same wrong assumption as the implementation. Always review test assertions critically — do they test the RIGHT behavior?

2 Coverage Analysis and Gap Filling

Once you have initial tests, coverage analysis reveals what's missing. Codex can read coverage reports and generate targeted tests for uncovered code.

Generate a coverage report:

> Run pytest with coverage and show me which lines in src/services/
> are not covered by tests.

Codex runs:

pytest --cov=src/services --cov-report=term-missing tests/

And interprets the output:

Name                          Stmts   Miss  Cover   Missing
-----------------------------------------------------------
src/services/auth.py             45     12    73%   34-38, 52-56, 71-74
src/services/payment.py          62     28    55%   23-35, 48-62, 78-90
src/services/notification.py     30      0   100%

Fill coverage gaps:

> Based on the coverage report, payment.py has the most gaps.
> Read payment.py and generate tests specifically for lines 23-35
> (the refund logic), lines 48-62 (the retry mechanism), and
> lines 78-90 (the webhook handler).

Coverage for error paths:

The most commonly missed coverage is error handling:

> Look at the uncovered lines in auth.py (34-38, 52-56, 71-74).
> These are likely error handling branches. Generate tests that
> trigger these error paths — invalid tokens, expired sessions,
> and database connection failures.

Targeting specific coverage goals:

> We need to get src/services/payment.py from 55% to 90% coverage.
> Generate the minimum number of tests needed to cover the remaining
> branches. Focus on the conditional logic (if/else branches).

Branch coverage vs line coverage:

> Run coverage with branch analysis and identify conditional
> branches that are only tested in one direction.
pytest --cov=src --cov-branch --cov-report=term-missing
NOTE
How It Works
Codex reads coverage reports and maps uncovered lines back to the source code. It then identifies what inputs would trigger those code paths and generates tests accordingly.
TIP
Tip
Don't aim for 100% coverage blindly. Some lines (platform-specific code, truly exceptional error paths) may not be worth testing. Focus coverage efforts on business logic and complex conditionals.

3 TDD Workflow with Codex

Test-Driven Development (TDD) means writing tests BEFORE the implementation. Codex makes this faster because it can generate comprehensive tests from a specification.

The TDD cycle with Codex:

1. Describe what the function should do
2. Ask Codex to write tests (they'll fail — no implementation yet)
3. Run tests to confirm they fail
4. Ask Codex to implement the function
5. Run tests to confirm they pass
6. Refactor if needed

Step 1: Describe the requirement

> I need a function called `parse_duration` in src/utils/time.py
> that converts human-readable duration strings to seconds.
> Examples:
> - "5m" → 300
> - "2h" → 7200
> - "1d" → 86400
> - "1h30m" → 5400
> - "500ms" → 0.5
> It should raise ValueError for invalid formats.

Step 2: Generate tests first

> Write comprehensive tests for this function in tests/test_time.py.
> Include all the examples above plus edge cases. Don't create the
> implementation yet.

Step 3: Confirm tests fail

> Run the tests — they should fail with ImportError since
> parse_duration doesn't exist yet.

Output: ImportError: cannot import name 'parse_duration' from 'src.utils.time'

Step 4: Implement to pass tests

> Now implement parse_duration in src/utils/time.py. Make all the
> tests pass.

Step 5: Run tests

> Run the tests again. Do they all pass?

Step 6: Refactor

> The implementation works but the regex is complex. Refactor it
> to be more readable while keeping all tests passing.

Why TDD with Codex works so well:

  • Codex writes thorough tests faster than you could manually
  • The tests act as a specification for the implementation
  • You get a complete test suite from the start (not an afterthought)
  • The implementation is guaranteed to match the specification
TIP
Tip
When doing TDD with Codex, describe the function behavior in detail but don't prescribe the implementation. Let the tests define WHAT it does and let Codex decide HOW.

4 Fixing Failing Tests

When tests fail, the question is: is the test wrong, or is the implementation wrong? Codex can help diagnose this.

Diagnosing test failures:

> Run pytest and three tests are failing. For each failure, tell me:
> 1. Is the test expectation correct?
> 2. Is the implementation wrong?
> 3. What's the fix?

Common failure patterns:

Pattern A — Test assumption is wrong:

> test_get_user_by_email asserts that the response status is 200,
> but our API returns 404 when the user is not in the database.
> The test isn't seeding a test user. Fix the test to set up proper
> test data first.

Pattern B — Implementation bug:

> test_calculate_discount asserts that a 20% discount on $100 gives
> $80, but the function returns $20 (it's returning the discount
> amount instead of the discounted price). Fix the implementation.

Pattern C — Environment issue:

> test_send_email fails with ConnectionRefusedError. This test needs
> the email sending to be mocked. Add a mock for the SMTP connection.

Batch fixing:

> Run the full test suite. Fix all failures, but explain each fix
> before making it so I can verify the approach.

Codex will:

  1. Run the tests
  2. For each failure, explain the root cause
  3. Propose a fix (wait for approval in untrusted policy)
  4. Move to the next failure

When tests fail after a refactor:

> I refactored the User model to use a separate Address object
> instead of flat address fields. Now 8 tests fail because they
> use the old field structure. Update the tests to use the new
> Address object format.
WARNING
Watch Out
When fixing failing tests, always verify that the fix is correct — not just that it makes the test pass. A test that asserts `True is True` will pass but is useless. Check that assertions still validate meaningful behavior.
NOTE
How It Works
Codex reads both the test file and the implementation to understand intent. It can usually determine whether the test or the code is 'wrong' based on function names, docstrings, and context.

5 Advanced Test Patterns

Beyond basic unit tests, Codex can generate more sophisticated test patterns including mocks, fixtures, parameterized tests, and integration tests.

Parameterized tests:

> Convert the individual test cases in test_validators.py into
> parameterized tests using pytest.mark.parametrize. Group related
> test data into clear tables.
@pytest.mark.parametrize("email,expected", [
    ("[email protected]", True),
    ("[email protected]", True),
    ("[email protected]", True),
    ("", False),
    ("no-at-sign", False),
    ("@no-local.com", False),
    ("spaces [email protected]", False),
])
def test_validate_email(email, expected):
    assert validate_email(email) == expected

Mock-based tests:

> Generate tests for src/services/payment.py. The PaymentService
> calls an external Stripe API — mock all external calls using
> unittest.mock.patch. Test both successful and failed payment
> scenarios.

Test fixtures:

> Create a conftest.py for the tests/ directory with reusable
> fixtures: a test database connection, a sample user, a sample
> order with items, and an authenticated client.

Integration tests:

> Write integration tests for the /api/orders endpoint that:
> 1. Create a test user and authenticate
> 2. Create an order via POST
> 3. Verify it appears in GET /api/orders
> 4. Update the order status via PATCH
> 5. Delete the order via DELETE
> 6. Verify it's gone
>
> Use a test database, not mocks.

Snapshot tests:

> Generate snapshot tests for the API response format. When the
> response structure changes unexpectedly, the test should fail.
> Use jest's toMatchSnapshot() for the JSON responses.

Performance tests:

> Add a basic performance test for the search function: it should
> return results within 100ms for a dataset of 10,000 items.
> Use pytest-benchmark or a simple timing assertion.
TIP
Tip
Tell Codex what testing pattern to use. Without guidance, it defaults to simple assertions. Explicitly ask for 'parameterized tests', 'mock-based tests', or 'integration tests' to get the right pattern.
TIP
Tip
After generating tests, review the assertions carefully. Codex sometimes generates tests that only check 'no exception was thrown' rather than verifying correct output. Good tests assert specific values.

Questions & Answers

Q: Can Codex-generated tests be trusted without review?
No. Always review generated tests, particularly the assertions. Common issues: tests that pass trivially (assert True), tests with wrong expected values (copied from a buggy implementation), and tests that test implementation details rather than behavior. The tests are a starting point, not a final product.
Q: How does Codex decide what edge cases to test?
Codex identifies edge cases from patterns: empty inputs, None/null values, boundary numbers (0, -1, MAX_INT), special characters in strings, empty lists, single-element lists, duplicate values, and type mismatches. You can also explicitly list edge cases you want tested in your prompt.
Q: Should I generate tests for third-party library code?
Generally no — test YOUR code that uses the library, not the library itself. Focus on testing your business logic, your integration points with the library, and your error handling around library calls. Mock the library in tests to isolate your code.
Q: How do I handle tests that need a database or external service?
Ask Codex to generate tests with mocks for unit testing, and separate integration tests that use a test database. For integration tests, have Codex create fixtures that set up and tear down test data. Use environment variables to switch between test and production databases.

Key Takeaways

  1. Generate then verify: Codex produces tests quickly, but always review assertions for correctness
  2. Coverage-driven testing: Use coverage reports to identify gaps and ask Codex to fill them specifically
  3. TDD is natural with Codex: Write tests from a spec, confirm they fail, then implement to pass
  4. Diagnose before fixing: When tests fail, determine whether the test or implementation is wrong
  5. Use advanced patterns: Ask explicitly for parameterized tests, mocks, and fixtures — don't settle for basic assertions
  6. Tests are documentation: Good test names and structure serve as living documentation of expected behavior

Next Steps: In Lesson 8 — Configuration & Customization, you'll learn how to configure Codex CLI to match your preferences and workflow.