File Operations
Learning Outcomes
- Create new files and directories using natural language instructions
- Read and summarize existing files without making changes
- Edit specific sections of files with precise instructions
- Perform bulk file operations across multiple files efficiently
- Interpret the diff view to understand proposed changes before approving
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | File operations as the core of coding with Codex |
| Demo 1 | 7 min | Creating files and directories |
| Demo 2 | 6 min | Reading and summarizing existing code |
| Demo 3 | 8 min | Editing files — targeted changes |
| Demo 4 | 7 min | Bulk operations across multiple files |
| Explain | 5 min | Reading diffs like a pro |
| Wrap-up | 4 min | Best practices for file operations |
Before You Begin
Pre-work:
- Complete Lesson 2 — Sandboxed Execution
- Create a test project with a few Python or JavaScript files
- Initialize git in your test project (
git init && git add . && git commit -m "initial")
Shopping List:
- Codex CLI installed and configured
- A test project with 3-5 existing source files
- Git initialized with a clean commit (so you can revert)
- Familiarity with basic diff format (+ and - lines)
The most common file operation is creation. Codex can generate files from scratch based on your description, including proper structure, imports, and boilerplate.
Creating a single file:
> Create a Python file called validators.py with functions to validate
> email addresses and phone numbers using regex
Codex will generate the complete file and show it for approval:
╭─ Creating file: validators.py ─────────────────╮
│ import re │
│ │
│ def validate_email(email: str) -> bool: │
│ """Validate an email address format.""" │
│ pattern = r'^[\w\.-]+@[\w\.-]+\.\w+$' │
│ return bool(re.match(pattern, email)) │
│ │
│ def validate_phone(phone: str) -> bool: │
│ """Validate a phone number format.""" │
│ pattern = r'^\+?1?\d{9,15}$' │
│ return bool(re.match(pattern, phone)) │
╰─────────────────────────────────────────────────╯
Creating directories with files:
> Create a directory structure for a REST API:
> - src/routes/ with an index.js
> - src/middleware/ with auth.js
> - src/models/ with user.js
> Each file should have a basic export skeleton
Codex will propose creating multiple files in sequence. In untrusted policy, you'll approve each one individually. In on-request policy, they'll all be created automatically.
Creating from examples:
You can reference existing files as patterns:
> Look at src/controllers/userController.js and create a similar
> controller for products called productController.js
Creating with context:
> Create a test file for validators.py that tests both functions
> with valid and invalid inputs. Use pytest.
Codex reads the existing validators.py to understand the functions, then generates appropriate tests.
Reading files is a zero-risk operation — Codex can always read files in your project without approval (in all modes). Use this to understand code before modifying it.
Summarize a file:
> What does src/auth.py do? Summarize its main functions and purpose.
Codex reads the file and provides a plain-English summary without making any changes.
Analyze patterns across files:
> What error handling pattern is used across the files in src/routes/?
> Are there any inconsistencies?
Codex reads all the files in the directory and reports its findings.
Find specific content:
> Which files import the database module? List them all.
> Find all TODO comments in this project
Understand relationships:
> Trace the request flow from the login endpoint to the database query.
> Which files are involved?
Comparing files:
> Compare userController.js and productController.js.
> What's different in their approach to error handling?
Reading with purpose:
The best read operations have a clear goal. Instead of "show me main.py", try:
- "What arguments does the CLI accept in main.py?"
- "Is there any input validation in form_handler.py?"
- "What's the return type of the process_data function?"
Editing existing files is where Codex truly shines. You describe what you want changed, and Codex produces a precise diff.
Surgical edits — change one thing:
> In config.py, change the database port from 5432 to 5433
Codex shows a minimal diff:
DATABASE_CONFIG = {
"host": "localhost",
- "port": 5432,
+ "port": 5433,
"name": "myapp",
}
Adding functionality:
> Add a logging statement at the beginning of each function in api.py
> that logs the function name and arguments
import logging
+
+ logger = logging.getLogger(__name__)
def get_users(limit=10):
+ logger.info(f"get_users called with limit={limit}")
return db.query("SELECT * FROM users LIMIT %s", limit)
def create_user(name, email):
+ logger.info(f"create_user called with name={name}, email={email}")
return db.execute("INSERT INTO users ...")
Refactoring:
> Refactor the calculate_total function to use a list comprehension
> instead of the for loop
Fixing issues:
> The parse_date function crashes when given an empty string.
> Add a guard clause at the top that returns None for empty input.
Being specific about location:
When files are large, be precise about where to make changes:
> In server.js, in the middleware section (around line 30-40),
> add CORS headers before the rate limiter
Codex can apply consistent changes across multiple files — a task that would take considerable time manually.
Rename a variable everywhere:
> Rename the variable 'usr' to 'current_user' in all Python files
> under src/
Codex will scan all matching files and propose edits for each one. In untrusted policy, you'll approve each file's changes individually.
Add a header to all files:
> Add a copyright header comment to the top of every .py file in this
> project: "# Copyright 2024 MyCompany. All rights reserved."
Consistent formatting changes:
> In all JavaScript files, change single quotes to double quotes
> for string literals
Pattern-based edits:
> Find all functions that catch generic Exception and change them
> to catch specific exceptions (ValueError, TypeError, etc.)
Updating imports after a refactor:
> I moved utils.py from src/ to src/helpers/. Update all import
> statements across the project to use the new path.
# For large bulk operations, on-request policy saves time
codex --approval-policy on-request "Add type hints to all functions in src/"
# For large bulk operations, on-request policy saves time
codex --approval-policy on-request "Add type hints to all functions in src/"
Verifying bulk changes:
After a bulk operation, verify the results:
> Check all the files you just modified. Did any of them have
> syntax errors introduced?
Or use git to review:
git diff --stat # See which files changed
git diff # See all changes in detail
Every file modification in Codex CLI is presented as a diff. Understanding how to read diffs quickly is essential for efficient approvals.
Diff anatomy:
╭─ Editing file: src/utils.py ───────────────────╮
│ │
│ @@ -12,8 +12,11 @@ │
│ │
│ def process_data(items): │
│ - result = [] │
│ - for item in items: │
│ - result.append(item.strip()) │
│ - return result │
│ + """Process a list of items. │
│ + │
│ + Args: │
│ + items: List of strings to process. │
│ + │
│ + Returns: │
│ + List of stripped strings. │
│ + """ │
│ + return [item.strip() for item in items] │
│ │
╰─────────────────────────────────────────────────╯
What each part means:
| Element | Meaning |
|---|---|
@@ -12,8 +12,11 @@ |
Location: starts at line 12, old version had 8 lines, new has 11 |
| Lines with no prefix | Unchanged context lines |
Lines starting with - |
Being removed (red in color terminals) |
Lines starting with + |
Being added (green in color terminals) |
Quick review checklist:
- Scope — Is the diff only touching what you asked for?
- Additions — Do the
+lines make sense for your request? - Removals — Are the
-lines things that should be removed? - Context — Do the unchanged lines confirm the right location?
- Side effects — Are there unexpected changes elsewhere in the file?
Multi-hunk diffs:
Sometimes a single file edit has multiple changed sections (hunks):
@@ -5,3 +5,4 @@
import os
+ import logging
@@ -20,4 +21,6 @@
def connect():
+ logging.info("Connecting to database")
conn = psycopg2.connect(...)
+ logging.info("Connected successfully")
return conn
Each @@ marker indicates a separate change location within the same file. Review each hunk independently.
Questions & Answers
Key Takeaways
- Create with context: Give Codex detailed descriptions and reference existing files as patterns for better results
- Read before editing: Use read operations to understand code before requesting changes
- Be precise with edits: Specific instructions produce focused diffs that are easy to review
- Bulk operations save time: Use on-request policy for repetitive changes across many files
- Master diff reading: Check scope, additions, removals, and side effects before approving
- Git is essential: Always have a clean commit before file operations so you can revert
Next Steps: In Lesson 4 — Multi-Step Tasks, you'll learn how to chain operations together for complex, multi-stage tasks.