Chain-of-Thought Prompting
Learning Outcomes
- Explain what chain-of-thought prompting is and why it improves output quality
- Identify when step-by-step reasoning helps and when it's unnecessary
- Use structured thinking prompts for architecture decisions
- Apply chain-of-thought techniques to debug complex issues
- Know when to request explicit reasoning vs letting the AI decide its approach
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | What is chain-of-thought? |
| Explain | 8 min | The science behind step-by-step reasoning |
| Demo | 10 min | CoT for architecture decisions |
| Demo | 8 min | CoT for debugging |
| Explain | 6 min | When CoT helps vs when it doesn't |
| Demo | 7 min | Structured thinking templates |
| Wrap-up | 3 min | Key takeaways |
Before You Begin
Pre-work:
- Complete Lesson 6 on System Prompts & Rules Files
- Have a project with at least one complex decision to make
- Think of a recent bug that took you time to solve
Shopping List:
- An AI coding tool (Claude Code, Cursor, Codex, or similar)
- A codebase with some complexity (multiple files, dependencies)
- A recent architectural question or bug to use as practice material
- Notepad for comparing AI outputs with and without CoT
Chain-of-thought (CoT) prompting is the technique of asking an AI to show its reasoning process before giving a final answer. Instead of jumping straight to a solution, the AI works through the problem step by step — and this process often produces better results.
Without CoT:
> What's the best database for my app?
Answer: Use PostgreSQL.
With CoT:
> Think through the database choice for my app step by step.
> Consider my requirements: real-time updates, complex queries,
> 10K concurrent users, and a small team.
Reasoning:
1. Real-time updates → need change streams or pub/sub capability
2. Complex queries → need strong query language, joins, aggregations
3. 10K concurrent → need connection pooling, horizontal read scaling
4. Small team → need managed options, simple operations
Analysis:
- MongoDB: good real-time (change streams), weak complex queries
- PostgreSQL: excellent queries, real-time via LISTEN/NOTIFY + logical replication
- DynamoDB: scales well but complex queries are painful
Recommendation: PostgreSQL with a connection pooler (PgBouncer)
and LISTEN/NOTIFY for real-time. The complex query requirement
rules out document stores, and the small team benefits from
PostgreSQL's mature ecosystem of managed services.
Why does this work?
When AI reasons step by step, it:
- Considers multiple factors before committing to an answer
- Catches contradictions in its own logic
- Gives you visibility into its reasoning (so you can spot errors)
- Produces more nuanced, context-appropriate answers
The key phrases that trigger CoT:
- "Think step by step"
- "Walk me through your reasoning"
- "Before answering, consider..."
- "Analyse the tradeoffs"
- "What factors should we weigh?"
- "Break this problem down"
Architecture decisions are where CoT shines brightest. These are multi-factor problems where the "right" answer depends on context.
Template for architectural CoT:
> I need to decide [architectural question].
> Before recommending a solution, think through:
> 1. What are the key requirements driving this decision?
> 2. What are the realistic options?
> 3. What are the tradeoffs of each option?
> 4. Given our constraints [list them], which option fits best?
> 5. What are the risks of this choice and how do we mitigate them?
Example: Choosing an authentication strategy
> We need to add authentication to our API. Think step by step:
>
> Context:
> - Multi-tenant SaaS app
> - Need to support SSO for enterprise customers
> - Currently 5 microservices, planning for 15
> - Team has experience with Express but not OAuth libraries
>
> Walk through the options and recommend an approach.
The AI will typically produce structured reasoning like:
Step 1: Requirements analysis
- Multi-tenant → token must carry tenant context
- SSO → must support SAML/OIDC federation
- Microservices → token validation must not require central call
- Team experience → simpler is better for initial implementation
Step 2: Options
A) Self-managed JWT with custom auth service
B) Auth0 / external identity provider
C) Keycloak (self-hosted)
Step 3: Tradeoff analysis
[Detailed comparison on each requirement]
Step 4: Recommendation with reasoning
[Specific choice tied back to the requirements]
Example: Choosing between patterns
> I'm implementing a notification system. Users can receive
> notifications via email, push, and in-app.
>
> Think step by step about whether I should use:
> A) An event bus (pub/sub) with subscribers per channel
> B) A strategy pattern with a notification dispatcher
> C) A simple switch/case in the notification service
>
> Consider: testability, adding new channels later, error handling
> if one channel fails, and team familiarity (we use Express,
> no message queues currently set up).
Getting deeper analysis:
If the first answer is too shallow, push further:
> You recommended option B. Now think through the implementation:
> - What interfaces do we need?
> - How do we handle partial failures (email sends but push fails)?
> - How does this look when we add SMS in 3 months?
> - Show me the key types/interfaces before writing code.
Debugging is detective work, and CoT helps the AI investigate systematically rather than guessing.
The debugging CoT template:
> I'm seeing [symptom]. Help me debug this step by step:
>
> 1. What are the possible causes of this symptom?
> 2. How can we narrow down which cause it is?
> 3. What should I check first (most likely cause)?
> 4. What evidence would confirm or rule out each theory?
Example: API returning stale data
> My API endpoint /users/:id sometimes returns stale data after
> an update. The update endpoint returns success, but a subsequent
> GET sometimes shows the old data.
>
> Think through possible causes step by step. Our stack:
> - Express API with PostgreSQL
> - Redis cache with 5-minute TTL
> - Read replicas for GET requests
> - Nginx reverse proxy with caching headers
The AI reasons through:
Possible causes (ordered by likelihood):
1. Redis cache not invalidated on update → most common cause
2. Read replica lag → PostgreSQL replication delay
3. Nginx caching the GET response → check Cache-Control headers
4. Client-side caching → browser/CDN caching the response
5. Race condition → GET arrives before UPDATE commits
Diagnostic steps:
1. Check: Does the problem resolve after 5 minutes? → Redis TTL
2. Check: Bypass cache and query primary directly → replica lag
3. Check: curl with no-cache headers → Nginx/client cache
4. Check: Does it only happen under load? → race condition
Most likely: #1 (Redis cache). The update endpoint probably
doesn't invalidate the cached user object.
Example: Intermittent test failure
> This test passes 90% of the time but occasionally fails with
> "expected 3 items but received 2". Think through why a test
> might be flaky in this way:
>
> The test creates 3 records, then queries for all records
> and asserts the count.
Deepening the investigation:
> You identified that another test might be cleaning up records.
> Think through: how do I confirm this? What would I look for
> in the test output? How do I fix it without slowing down the
> test suite?
Pattern: "What would a senior engineer check?"
> A senior engineer is debugging this memory leak.
> Walk through their thought process:
> - What metrics would they look at first?
> - What tools would they use?
> - What would they test and in what order?
> - When would they escalate vs keep investigating?
Chain-of-thought is powerful, but it's not always the right tool. Here's a guide:
CoT is most valuable for:
| Situation | Why CoT Helps |
|---|---|
| Multi-factor decisions | Forces consideration of all factors |
| Debugging complex issues | Systematic investigation vs random guessing |
| Architecture choices | Tradeoff analysis needs explicit reasoning |
| Code review | Explaining why something is problematic |
| Input validation | Tracing every place external data enters needs step-by-step reasoning |
| Performance optimization | Need to identify bottlenecks before optimising |
CoT is unnecessary for:
| Situation | Why Skip CoT |
|---|---|
| Simple code generation | "Write a function to sort an array" — just do it |
| Formatting/styling | "Convert these to camelCase" — no reasoning needed |
| Boilerplate generation | "Create a React component" — pattern is straightforward |
| Looking up facts | "What's the Express middleware signature?" — just answer |
| Small, clear tasks | "Add a null check before line 15" — obvious intent |
CoT can actually hurt when:
- The task is simple and CoT adds unnecessary verbosity
- You need quick iteration and the reasoning slows output
- The problem is well-defined with one obvious solution
- You're generating large amounts of code and don't need justification for each line
Choosing the right level of reasoning:
No CoT: "Write a function that validates email addresses"
Light CoT: "Write an email validator and briefly explain your regex choice"
Medium CoT: "Design an email validation strategy — consider edge cases before implementing"
Heavy CoT: "Think through email validation thoroughly: what standards exist,
what tradeoffs between strict/permissive, what are the edge cases,
then implement the best approach for a user registration form"
Rule of thumb: If you'd want a colleague to explain their reasoning before implementing, use CoT. If you'd just say "go do it", skip CoT.
Here are reusable CoT templates for common development scenarios. Adapt these to your specific needs.
Template: Technology Evaluation
> Evaluate [technology] for our use case. Think through:
> 1. What problem does it solve that we currently have?
> 2. What's the learning curve for our team?
> 3. What's the operational burden (hosting, monitoring, upgrades)?
> 4. What are we giving up by adopting it?
> 5. What's the exit strategy if it doesn't work out?
> 6. Based on this analysis, should we adopt it? For what scope?
Template: Code Design Decision
> Before implementing, think through the design:
> 1. What are the inputs and outputs of this system?
> 2. What are the key operations and their frequency?
> 3. What are the failure modes and how should each be handled?
> 4. What will change most often in the future?
> 5. Design the interfaces first, then describe the implementation.
Template: Bug Root Cause Analysis
> Help me find the root cause. Think step by step:
> 1. What is the exact symptom? (What happens vs what should happen)
> 2. When did it start? (What changed recently?)
> 3. What are all possible causes? (List at least 5)
> 4. Rank causes by likelihood given the evidence
> 5. For the top cause, what confirms or denies it?
> 6. Suggest a fix and explain why it addresses the root cause
Template: Performance Investigation
> This operation is slow ([current time] vs [target time]).
> Think through the performance analysis:
> 1. What are all the operations happening in this path?
> 2. Which operations are likely bottlenecks? (I/O, computation, network)
> 3. What would a flame graph likely show?
> 4. What's the quickest win vs the most impactful fix?
> 5. Are there caching opportunities? At what layer?
Template: Input Handling Review
> Review how this code handles external input. Think through:
> 1. Where does data from outside the code enter?
> 2. What unexpected input could it receive? (empty, huge, wrong type, wrong encoding)
> 3. Is every query parameterised, and every value validated before use?
> 4. What breaks, and how far does the damage spread, if one input is bad?
> 5. Where would a second check catch what the first one misses?
Combining templates in practice:
You can chain templates for thorough analysis:
> First, use the bug root cause template to identify the issue.
> Then, use the code design template to design a proper fix.
> Finally, think through whether the fix introduces any new input-handling problems.
Questions & Answers
Key Takeaways
- CoT improves complex decisions — step-by-step reasoning catches flaws that direct answers miss
- Be specific about what to think through — "analyse tradeoffs of X vs Y" beats "think step by step"
- Use CoT for debugging — systematic investigation beats random guessing
- Skip CoT for simple tasks — it adds verbosity without value for straightforward requests
- Build templates — reusable CoT structures ensure thorough analysis every time
- Verify the reasoning — CoT lets you check the logic, not just the conclusion
Next Steps: In Lesson 8 — Multi-Turn Conversations, you'll learn how to build effective context across multiple messages for larger, more complex tasks.