Advanced Automation

60 min advanced Lesson 10

Learning Outcomes

  • Write shell scripts that orchestrate Codex CLI for batch file processing
  • Build automated code review pipelines that run on multiple repositories
  • Create multi-stage processing workflows combining Codex with other CLI tools
  • Implement guardrails and quality checks for fully automated Codex operations
  • Design scalable automation patterns that handle failures gracefully

Lesson Plan

Segment Duration Topic
Intro 4 min From interactive to automated — when and why
Explain 6 min Scripting fundamentals with Codex CLI
Demo 1 10 min Batch processing — migrating 50 files
Demo 2 10 min Automated review pipeline
Demo 3 8 min Multi-tool orchestration (Codex + lint + test)
Explain 6 min Error handling and retry strategies
Demo 4 8 min Building a complete automation pipeline
Demo 5 5 min Monitoring and logging automated runs
Wrap-up 3 min Automation maturity model

Before You Begin

Pre-work:

  • Complete Lesson 9 — Integration Workflows
  • Comfortable with shell scripting (bash/zsh)
  • Have a multi-file project that needs repetitive changes
  • Understand API rate limits and cost implications

Shopping List:

  • Codex CLI installed and configured
  • Shell scripting environment (bash or zsh)
  • A project with 10+ files for batch processing exercises
  • jq installed (for JSON processing in scripts)
  • Understanding of exit codes and error handling in shell scripts

1 Scripting Fundamentals

Automating Codex means running it non-interactively with predictable inputs and parseable outputs. Understanding the CLI flags for automation is essential.

Key flags for automation:

Flag Purpose Example
--quiet Suppress UI, output result only codex --quiet "explain main.py"
--approval-policy never No approval prompts Unattended execution
--model Specify model Cost/speed control per task

Basic automation pattern:

#!/bin/bash
# scripts/codex-task.sh — Basic automation template

set -euo pipefail  # Exit on error, undefined vars, pipe failures

PROJECT_DIR="/path/to/project"
LOG_FILE="codex-automation.log"

# Ensure we're in the right directory
cd "$PROJECT_DIR"

# Ensure clean git state
if [ -n "$(git status --porcelain)" ]; then
  echo "ERROR: Working directory not clean. Commit or stash changes first."
  exit 1
fi

# Run Codex task
echo "[$(date)] Starting Codex automation..." | tee -a "$LOG_FILE"

OUTPUT=$(codex --quiet --approval-policy never \
  "Add input validation to all route handlers in src/routes/")

echo "$OUTPUT" | tee -a "$LOG_FILE"

# Verify result
if [ -n "$(git status --porcelain)" ]; then
  echo "[$(date)] Changes detected. Running tests..." | tee -a "$LOG_FILE"
  npm test 2>&1 | tee -a "$LOG_FILE"

  if [ $? -eq 0 ]; then
    echo "[$(date)] Tests passed. Committing." | tee -a "$LOG_FILE"
    git add -A
    git commit -m "feat: add input validation to route handlers (automated)"
  else
    echo "[$(date)] Tests failed. Reverting." | tee -a "$LOG_FILE"
    git checkout -- .
    exit 1
  fi
else
  echo "[$(date)] No changes made." | tee -a "$LOG_FILE"
fi
#!/bin/bash
# scripts/codex-task.sh — Basic automation template (Git Bash)

set -euo pipefail

PROJECT_DIR="/c/projects/myproject"
LOG_FILE="codex-automation.log"

cd "$PROJECT_DIR"

# Ensure clean git state
if [ -n "$(git status --porcelain)" ]; then
  echo "ERROR: Working directory not clean. Commit or stash changes first."
  exit 1
fi

# Run Codex task
echo "[$(date)] Starting Codex automation..." | tee -a "$LOG_FILE"

OUTPUT=$(codex --quiet --approval-policy never \
  "Add input validation to all route handlers in src/routes/")

echo "$OUTPUT" | tee -a "$LOG_FILE"

# Verify result
if [ -n "$(git status --porcelain)" ]; then
  echo "[$(date)] Changes detected. Running tests..." | tee -a "$LOG_FILE"
  npm test 2>&1 | tee -a "$LOG_FILE"

  if [ $? -eq 0 ]; then
    echo "[$(date)] Tests passed. Committing." | tee -a "$LOG_FILE"
    git add -A
    git commit -m "feat: add input validation to route handlers (automated)"
  else
    echo "[$(date)] Tests failed. Reverting." | tee -a "$LOG_FILE"
    git checkout -- .
    exit 1
  fi
else
  echo "[$(date)] No changes made." | tee -a "$LOG_FILE"
fi

Key principles for automation:

  1. Always check git status before running — require a clean state
  2. Always run tests after changes — revert if they fail
  3. Log everything — you need a record of what happened
  4. Use never policy — automation can't answer approval prompts
  5. Set exit codes — so calling scripts know if it succeeded
WARNING
Watch Out
Full-auto mode in scripts means no human review before execution. Always have automated checks (tests, linting) as your quality gate. Never run full-auto automation on production branches without CI/CD guardrails.
NOTE
How It Works
When Codex runs with --quiet and --approval-policy never, it behaves like a pure function: input (prompt + files) → output (file changes + stdout). This makes it predictable enough for scripting.

2 Batch Processing

Batch processing applies the same transformation to many files. This is ideal for migrations, style changes, or adding consistent patterns.

Example: Adding type hints to all Python files

#!/bin/bash
# scripts/batch-type-hints.sh

set -euo pipefail

RESULTS_LOG="batch-results.log"
FAILED_FILES=()
SUCCESS_COUNT=0

# Find all Python files without type hints
FILES=$(find src/ -name "*.py" -type f)

echo "Found $(echo "$FILES" | wc -l) Python files to process"
echo "---" > "$RESULTS_LOG"

for file in $FILES; do
  echo "Processing: $file"

  # Create a checkpoint
  git stash push -m "pre-batch-$file" --quiet 2>/dev/null || true

  # Run Codex on this file
  codex --quiet --approval-policy never \
    "Add type hints to all functions in $file. Don't change any logic, only add type annotations for parameters and return values." \
    2>> "$RESULTS_LOG"

  # Verify the file is still valid Python
  if python -c "import ast; ast.parse(open('$file').read())" 2>/dev/null; then
    SUCCESS_COUNT=$((SUCCESS_COUNT + 1))
    echo "  ✓ Success" | tee -a "$RESULTS_LOG"
  else
    echo "  ✗ Syntax error — reverting" | tee -a "$RESULTS_LOG"
    git checkout -- "$file"
    FAILED_FILES+=("$file")
  fi
done

echo ""
echo "=== Batch Complete ==="
echo "Successful: $SUCCESS_COUNT"
echo "Failed: $FAILED_FILES_COUNT"

if [ $FAILED_FILES_COUNT -gt 0 ]; then
  echo "Failed files:"
  printf '  %s\n' "${FAILED_FILES[@]}"
fi

# Run full test suite
echo ""
echo "Running tests..."
if pytest tests/ -q; then
  echo "All tests pass. Safe to commit."
  git add -A
  git commit -m "refactor: add type hints to src/ (batch automated)"
else
  echo "Tests failed! Review changes manually."
  exit 1
fi

Example: Migrating API response format

#!/bin/bash
# scripts/batch-api-migration.sh
# Migrate all endpoints from { data: ... } to { result: ..., meta: {} }

set -euo pipefail

ROUTES_DIR="src/routes"

for route_file in "$ROUTES_DIR"/*.js; do
  echo "Migrating: $route_file"

  codex --quiet --approval-policy never \
    "In $route_file, change all API responses from the format
     res.json({ data: value }) to res.json({ result: value, meta: { timestamp: Date.now() } }).
     Don't change error responses. Only modify success responses."

  # Verify with eslint
  npx eslint "$route_file" --quiet || {
    echo "ESLint failed for $route_file — reverting"
    git checkout -- "$route_file"
  }
done

Rate limiting batch operations:

# Add a delay between files to avoid API rate limits
for file in $FILES; do
  codex --quiet --approval-policy never "Process $file..."
  sleep 2  # Pause 2 seconds between API calls
done

Parallel processing (with caution):

# Process files in parallel (max 3 at a time)
echo "$FILES" | xargs -P 3 -I {} bash -c '
  codex --quiet --approval-policy never "Add docstrings to {}"
'
TIP
Tip
Always process one file first manually to verify the transformation works correctly. Then batch the rest. This catches prompt issues early before burning through API tokens on 50 files.
WARNING
Watch Out
Parallel processing can cause race conditions if files depend on each other (shared imports, etc.). Use sequential processing for interdependent files.

3 Automated Review Pipelines

Build pipelines that automatically review code for issues, generate reports, and even fix problems without human intervention.

Review pipeline architecture:

┌─────────────┐    ┌─────────────┐    ┌─────────────┐    ┌─────────────┐
│  Detect     │ →  │  Analyze    │ →  │  Fix        │ →  │  Verify     │
│  Changes    │    │  with Codex │    │  with Codex │    │  with Tests │
└─────────────┘    └─────────────┘    └─────────────┘    └─────────────┘

Complete review pipeline script:

#!/bin/bash
# scripts/review-pipeline.sh
# Run on a feature branch to review, fix, and verify

set -euo pipefail

BRANCH=$(git branch --show-current)
BASE="main"
REPORT_FILE="review-report.md"

echo "# Code Review Report" > "$REPORT_FILE"
echo "Branch: $BRANCH" >> "$REPORT_FILE"
echo "Date: $(date)" >> "$REPORT_FILE"
echo "" >> "$REPORT_FILE"

# Stage 1: Identify what changed
echo "## Changed Files" >> "$REPORT_FILE"
CHANGED_FILES=$(git diff --name-only "$BASE"..."$BRANCH")
echo "$CHANGED_FILES" >> "$REPORT_FILE"
echo "" >> "$REPORT_FILE"

# Stage 2: AI Review
echo "## Review Findings" >> "$REPORT_FILE"
REVIEW=$(codex --quiet \
  "Review git diff $BASE...$BRANCH. Report:
   1. Potential bugs (with file and line)
   2. Security concerns
   3. Missing error handling
   4. Performance issues
   Be specific and actionable. No false positives.")

echo "$REVIEW" >> "$REPORT_FILE"
echo "" >> "$REPORT_FILE"

# Stage 3: Automated fixes for low-risk issues
echo "## Automated Fixes" >> "$REPORT_FILE"

# Fix missing error handling
codex --quiet --approval-policy never \
  "In the changed files ($CHANGED_FILES), add try-catch blocks
   around any async operations that lack error handling.
   Use the project's existing error handling pattern." \
  2>/dev/null

if [ -n "$(git status --porcelain)" ]; then
  echo "- Added error handling to async operations" >> "$REPORT_FILE"
fi

# Stage 4: Verify fixes
echo "" >> "$REPORT_FILE"
echo "## Verification" >> "$REPORT_FILE"

if npm test 2>/dev/null; then
  echo "- All tests pass after fixes" >> "$REPORT_FILE"
  git add -A
  git commit -m "fix: automated error handling improvements"
else
  echo "- Tests failed after fixes — reverting" >> "$REPORT_FILE"
  git checkout -- .
fi

echo ""
echo "Review complete. Report: $REPORT_FILE"
cat "$REPORT_FILE"

Multi-repository review:

#!/bin/bash
# scripts/review-all-repos.sh
# Review multiple repos in a monorepo or workspace

REPOS=("service-auth" "service-payments" "service-notifications")

for repo in "${REPOS[@]}"; do
  echo "=== Reviewing: $repo ==="
  cd "$repo"

  codex --quiet \
    "Review this service for:
     1. Outdated dependencies (check package.json)
     2. Unused exports (defined but never imported elsewhere)
     3. TODO/FIXME comments that should be addressed
     4. Functions longer than 50 lines that should be split
     Summarize findings concisely."

  cd ..
  echo ""
done

Scheduled review (cron job):

# Add to crontab: run every Monday at 9am
# crontab -e
# 0 9 * * 1 /path/to/scripts/weekly-review.sh

#!/bin/bash
# scripts/weekly-review.sh

cd /path/to/project

codex --quiet \
  "Perform a weekly code health check:
   1. Find any new TODO/FIXME comments added this week
   2. Check for files that have grown beyond 300 lines
   3. Identify any test files that haven't been updated in 30+ days
   4. Flag any hardcoded values that should be in config
   Output as a markdown report." > weekly-report.md

# Email the report
cat weekly-report.md | mail -s "Weekly Code Review" [email protected]
NOTE
How It Works
Review pipelines use Codex in read-only mode (just analysis) for Stage 2, then switch to write mode (full-auto) for Stage 3. This separation means the risky part (writing) is smaller and easier to verify.
TIP
Tip
Keep automated fixes conservative — only fix issues where the correct solution is unambiguous (missing error handling, unused imports, formatting). Leave complex logic changes for human review.

4 Multi-Tool Orchestration

The most powerful automation combines Codex with other tools in a pipeline. Each tool does what it's best at.

Pipeline: Codex + ESLint + Prettier + Tests

#!/bin/bash
# scripts/quality-pipeline.sh
# Complete quality pass: AI review → lint fix → format → test

set -euo pipefail

echo "=== Stage 1: AI Code Improvements ==="
codex --quiet --approval-policy never \
  "Review all files in src/ for:
   - Functions without JSDoc comments (add them)
   - Magic numbers (extract to named constants)
   - Repeated code blocks (extract to helper functions)
   Apply the fixes."

echo "=== Stage 2: Lint Fix ==="
npx eslint src/ --fix --quiet

echo "=== Stage 3: Format ==="
npx prettier --write "src/**/*.{ts,js}" --log-level error

echo "=== Stage 4: Type Check ==="
npx tsc --noEmit
if [ $? -ne 0 ]; then
  echo "Type errors detected. Asking Codex to fix..."
  codex --quiet --approval-policy never \
    "Run tsc --noEmit and fix all type errors reported."
fi

echo "=== Stage 5: Test ==="
npm test

echo "=== Stage 6: Coverage Check ==="
COVERAGE=$(npm test -- --coverage --coverageReporters=text-summary 2>&1 | grep "Statements" | awk '{print $3}')
echo "Coverage: $COVERAGE"

if (( $(echo "$COVERAGE < 80" | bc -l) )); then
  echo "Coverage below 80%. Generating additional tests..."
  codex --quiet --approval-policy never \
    "Run tests with coverage. Find uncovered files and generate tests to bring coverage above 80%."
  npm test
fi

echo "=== Pipeline Complete ==="
git add -A
git diff --cached --stat

Pipeline: Codex + Dependency Audit

#!/bin/bash
# scripts/dependency-pipeline.sh

echo "=== Running dependency audit ==="
AUDIT_OUTPUT=$(npm audit --json 2>/dev/null)

VULNERABILITIES=$(echo "$AUDIT_OUTPUT" | jq '.metadata.vulnerabilities.high + .metadata.vulnerabilities.critical')

if [ "$VULNERABILITIES" -gt 0 ]; then
  echo "Found $VULNERABILITIES high/critical vulnerabilities"

  # Ask Codex to analyze and suggest fixes
  codex --quiet \
    "npm audit found these vulnerabilities:
     $(echo "$AUDIT_OUTPUT" | jq '.vulnerabilities | to_entries[] | select(.value.severity == \"high\" or .value.severity == \"critical\") | .value.name + \": \" + .value.title' -r)

     For each one, tell me:
     1. Is there a safe upgrade path?
     2. If not, is there an alternative package?
     3. What code changes would be needed?"
fi

echo "=== Checking for secrets in code ==="
# Use grep to find potential secrets, then Codex to verify
POTENTIAL_SECRETS=$(grep -rn "password\|secret\|api_key\|token" src/ --include="*.ts" --include="*.js" | grep -v "test" | grep -v ".example")

if [ -n "$POTENTIAL_SECRETS" ]; then
  codex --quiet \
    "These lines might contain hardcoded secrets:
     $POTENTIAL_SECRETS

     For each one, tell me:
     - Is this actually a hardcoded secret? (vs a variable name or config key)
     - If it is, how should it be moved to environment variables?"
fi

Pipeline: Codex + Database Migration

#!/bin/bash
# scripts/migration-pipeline.sh

# Step 1: Codex generates the migration
codex --quiet --approval-policy never \
  "Create a database migration in migrations/ that adds an 'archived_at'
   timestamp column to the users table. Follow the pattern in the most
   recent migration file."

# Step 2: Run the migration against test DB
DATABASE_URL="postgresql://localhost/test_db" npx knex migrate:latest

# Step 3: Codex updates the model
codex --quiet --approval-policy never \
  "Update src/models/User.ts to include the new archived_at field.
   Add a soft-delete method that sets archived_at to the current time."

# Step 4: Generate tests for new functionality
codex --quiet --approval-policy never \
  "Write tests for the soft-delete functionality in User model."

# Step 5: Run all tests
npm test
TIP
Tip
Design pipelines where each stage validates the previous stage's output. If lint fails after Codex edits, that catches syntax issues. If tests fail after lint, that catches logic issues. Each stage is a safety net.
WARNING
Watch Out
Multi-tool pipelines can be slow and expensive. Each Codex call is an API request. Optimize by batching related changes into single prompts rather than calling Codex separately for each file.

5 Error Handling and Monitoring

Robust automation handles failures gracefully. API timeouts, rate limits, and unexpected outputs all need to be managed.

Retry logic for API failures:

#!/bin/bash
# scripts/lib/codex-with-retry.sh

codex_with_retry() {
  local max_attempts=3
  local attempt=1
  local delay=5
  local prompt="$1"

  while [ $attempt -le $max_attempts ]; do
    OUTPUT=$(codex --quiet --approval-policy never "$prompt" 2>&1)
    EXIT_CODE=$?

    if [ $EXIT_CODE -eq 0 ]; then
      echo "$OUTPUT"
      return 0
    fi

    # Check if it's a rate limit error
    if echo "$OUTPUT" | grep -qi "rate limit\|429"; then
      echo "Rate limited. Waiting ${delay}s before retry..." >&2
      sleep $delay
      delay=$((delay * 2))  # Exponential backoff
    elif echo "$OUTPUT" | grep -qi "timeout\|connection"; then
      echo "Connection error. Retrying in ${delay}s..." >&2
      sleep $delay
    else
      echo "Non-retryable error: $OUTPUT" >&2
      return 1
    fi

    attempt=$((attempt + 1))
  done

  echo "Failed after $max_attempts attempts" >&2
  return 1
}

# Usage:
# source scripts/lib/codex-with-retry.sh
# RESULT=$(codex_with_retry "Add error handling to src/api.js")

Comprehensive error handling in batch scripts:

#!/bin/bash
# scripts/safe-batch.sh

set -euo pipefail

ERRORS=()
SUCCESSES=()
SKIPPED=()

process_file() {
  local file="$1"

  # Skip files that are too large
  local lines=$(wc -l < "$file")
  if [ "$lines" -gt 500 ]; then
    SKIPPED+=("$file (too large: $lines lines)")
    return 0
  fi

  # Attempt processing
  if OUTPUT=$(codex_with_retry "Add type hints to $file" 2>/dev/null); then
    # Verify the file is still valid
    if python -c "import ast; ast.parse(open('$file').read())" 2>/dev/null; then
      SUCCESSES+=("$file")
    else
      git checkout -- "$file"
      ERRORS+=("$file (syntax error after modification)")
    fi
  else
    ERRORS+=("$file (Codex API failure)")
  fi
}

# Process all files
for file in src/**/*.py; do
  process_file "$file"
done

# Summary report
echo ""
echo "=== Batch Processing Summary ==="
echo "Successful: $SUCCESSES_COUNT"
echo "Failed:     $ERRORS_COUNT"
echo "Skipped:    $SKIPPED_COUNT"

if [ $ERRORS_COUNT -gt 0 ]; then
  echo ""
  echo "Failed files:"
  printf '  - %s\n' "${ERRORS[@]}"
fi

Monitoring and alerting:

#!/bin/bash
# scripts/monitored-automation.sh

LOG_DIR="logs/codex-automation"
mkdir -p "$LOG_DIR"

RUN_ID=$(date +%Y%m%d_%H%M%S)
LOG_FILE="$LOG_DIR/run_$RUN_ID.log"

# Redirect all output to log file AND terminal
exec > >(tee -a "$LOG_FILE") 2>&1

echo "Run ID: $RUN_ID"
echo "Started: $(date)"
echo ""

# Your automation here...
codex --quiet --approval-policy never "..." || {
  echo "ALERT: Automation failed at $(date)"
  # Send alert (email, team chat, PagerDuty, etc.)
  curl -X POST "$ALERT_WEBHOOK_URL" \
    -H 'Content-Type: application/json' \
    -d "{\"text\": \"Codex automation failed. Run: $RUN_ID. Check logs.\"}"
  exit 1
}

echo ""
echo "Completed: $(date)"

# Archive logs older than 30 days
find "$LOG_DIR" -name "*.log" -mtime +30 -delete

Cost tracking:

#!/bin/bash
# Track API usage per automation run

COST_LOG="logs/codex-costs.csv"

# Initialize CSV header if file doesn't exist
if [ ! -f "$COST_LOG" ]; then
  echo "date,script,model,duration_seconds,files_processed" > "$COST_LOG"
fi

START_TIME=$(date +%s)
# ... your automation ...
END_TIME=$(date +%s)
DURATION=$((END_TIME - START_TIME))

echo "$(date +%Y-%m-%d),$0,gpt-5.4-mini,$DURATION,$FILE_COUNT" >> "$COST_LOG"
NOTE
How It Works
API failures are inevitable in automation — network issues, rate limits, and service outages all occur. Retry logic with exponential backoff handles transient failures. Logging provides audit trails for debugging persistent issues.
TIP
Tip
Run all new automation scripts with a --dry-run flag first that shows what WOULD happen without actually calling the API. Add dry-run support to your scripts during development.

Questions & Answers

Q: How do I handle Codex producing different results each run?
AI outputs are non-deterministic — the same prompt may produce different code each time. For automation, this means: (1) always validate output with tests and linting, (2) don't rely on exact output format for parsing, (3) use specific prompts that constrain the solution space, and (4) accept that some runs may produce sub-optimal results that need re-running.
Q: What's the cost of running batch automation on a large codebase?
Costs depend on model, prompt size, and file count. A rough estimate: processing 50 files with gpt-5.4-mini costs approximately $0.50-2.00 total. With gpt-5.5, multiply by 5-10x. To control costs: use gpt-5.4-mini for repetitive tasks, batch related changes into fewer API calls, and set daily spending limits on your API key.
Q: Can I run Codex automation in Docker containers for isolation?
Yes, and this is recommended for full-auto automation. Run Codex inside a Docker container that contains only your project files. This adds OS-level isolation on top of Codex's own sandbox. Mount your project as a volume, set the API key as an environment variable, and run your automation script inside the container.
Q: How do I handle automation that needs to modify files across multiple repos?
Since Codex's sandbox limits it to one directory, handle multi-repo operations by: (1) cloning all repos into subdirectories of a parent folder, (2) running Codex from the parent with appropriate file paths, or (3) running separate Codex instances per repo in a loop. Option 3 is simpler and keeps sandboxes properly isolated.

Key Takeaways

  1. Script for repeatability: Convert manual Codex sessions into scripts for tasks you'll do more than once
  2. Batch with safeguards: Process files in loops but validate each result (syntax check, lint, tests) before moving on
  3. Pipeline design: Combine Codex with linters, formatters, type checkers, and test runners — each validates the previous stage
  4. Handle failures gracefully: Retry transient errors, revert bad changes, log everything, alert on persistent failures
  5. Start conservative: Run one file manually first, then small batches, then full automation — build confidence incrementally
  6. Monitor costs: Track API usage per script and set spending limits to prevent runaway costs

Next Steps: You've completed the Learn Codex course! Apply these automation patterns to your daily workflow, starting with simple scripts and gradually building more sophisticated pipelines as your confidence grows. Return to the course overview to review any lessons or explore supplemental materials.