Chain-of-Thought Prompting

45 min intermediate Lesson 7

Learning Outcomes

  • Explain what chain-of-thought prompting is and why it improves output quality
  • Identify when step-by-step reasoning helps and when it's unnecessary
  • Use structured thinking prompts for architecture decisions
  • Apply chain-of-thought techniques to debug complex issues
  • Know when to request explicit reasoning vs letting the AI decide its approach

Lesson Plan

Segment Duration Topic
Intro 3 min What is chain-of-thought?
Explain 8 min The science behind step-by-step reasoning
Demo 10 min CoT for architecture decisions
Demo 8 min CoT for debugging
Explain 6 min When CoT helps vs when it doesn't
Demo 7 min Structured thinking templates
Wrap-up 3 min Key takeaways

Before You Begin

Pre-work:

  • Complete Lesson 6 on System Prompts & Rules Files
  • Have a project with at least one complex decision to make
  • Think of a recent bug that took you time to solve

Shopping List:

  • An AI coding tool (Claude Code, Cursor, Codex, or similar)
  • A codebase with some complexity (multiple files, dependencies)
  • A recent architectural question or bug to use as practice material
  • Notepad for comparing AI outputs with and without CoT

1 What Chain-of-Thought Means

Chain-of-thought (CoT) prompting is the technique of asking an AI to show its reasoning process before giving a final answer. Instead of jumping straight to a solution, the AI works through the problem step by step — and this process often produces better results.

Without CoT:

> What's the best database for my app?

Answer: Use PostgreSQL.

With CoT:

> Think through the database choice for my app step by step.
> Consider my requirements: real-time updates, complex queries,
> 10K concurrent users, and a small team.

Reasoning:
1. Real-time updates → need change streams or pub/sub capability
2. Complex queries → need strong query language, joins, aggregations
3. 10K concurrent → need connection pooling, horizontal read scaling
4. Small team → need managed options, simple operations

Analysis:
- MongoDB: good real-time (change streams), weak complex queries
- PostgreSQL: excellent queries, real-time via LISTEN/NOTIFY + logical replication
- DynamoDB: scales well but complex queries are painful

Recommendation: PostgreSQL with a connection pooler (PgBouncer) 
and LISTEN/NOTIFY for real-time. The complex query requirement 
rules out document stores, and the small team benefits from 
PostgreSQL's mature ecosystem of managed services.

Why does this work?

When AI reasons step by step, it:

  • Considers multiple factors before committing to an answer
  • Catches contradictions in its own logic
  • Gives you visibility into its reasoning (so you can spot errors)
  • Produces more nuanced, context-appropriate answers

The key phrases that trigger CoT:

  • "Think step by step"
  • "Walk me through your reasoning"
  • "Before answering, consider..."
  • "Analyse the tradeoffs"
  • "What factors should we weigh?"
  • "Break this problem down"
NOTE
Key Insight
Chain-of-thought isn't just about getting a better answer — it's about getting a verifiable answer. When you can see the reasoning, you can check whether it makes sense for your specific situation.

2 CoT for Architecture Decisions

Architecture decisions are where CoT shines brightest. These are multi-factor problems where the "right" answer depends on context.

Template for architectural CoT:

> I need to decide [architectural question]. 

> Before recommending a solution, think through:
> 1. What are the key requirements driving this decision?
> 2. What are the realistic options?
> 3. What are the tradeoffs of each option?
> 4. Given our constraints [list them], which option fits best?
> 5. What are the risks of this choice and how do we mitigate them?

Example: Choosing an authentication strategy

> We need to add authentication to our API. Think step by step:
>
> Context:
> - Multi-tenant SaaS app
> - Need to support SSO for enterprise customers
> - Currently 5 microservices, planning for 15
> - Team has experience with Express but not OAuth libraries
>
> Walk through the options and recommend an approach.

The AI will typically produce structured reasoning like:

Step 1: Requirements analysis
- Multi-tenant → token must carry tenant context
- SSO → must support SAML/OIDC federation
- Microservices → token validation must not require central call
- Team experience → simpler is better for initial implementation

Step 2: Options
A) Self-managed JWT with custom auth service
B) Auth0 / external identity provider
C) Keycloak (self-hosted)

Step 3: Tradeoff analysis
[Detailed comparison on each requirement]

Step 4: Recommendation with reasoning
[Specific choice tied back to the requirements]

Example: Choosing between patterns

> I'm implementing a notification system. Users can receive 
> notifications via email, push, and in-app.
>
> Think step by step about whether I should use:
> A) An event bus (pub/sub) with subscribers per channel
> B) A strategy pattern with a notification dispatcher
> C) A simple switch/case in the notification service
>
> Consider: testability, adding new channels later, error handling 
> if one channel fails, and team familiarity (we use Express, 
> no message queues currently set up).

Getting deeper analysis:

If the first answer is too shallow, push further:

> You recommended option B. Now think through the implementation:
> - What interfaces do we need?
> - How do we handle partial failures (email sends but push fails)?
> - How does this look when we add SMS in 3 months?
> - Show me the key types/interfaces before writing code.
TIP
Tip
Save your best architectural CoT prompts as templates. You'll face similar decisions across projects — having a proven prompt structure saves time and ensures thorough analysis.

3 CoT for Debugging

Debugging is detective work, and CoT helps the AI investigate systematically rather than guessing.

The debugging CoT template:

> I'm seeing [symptom]. Help me debug this step by step:
>
> 1. What are the possible causes of this symptom?
> 2. How can we narrow down which cause it is?
> 3. What should I check first (most likely cause)?
> 4. What evidence would confirm or rule out each theory?

Example: API returning stale data

> My API endpoint /users/:id sometimes returns stale data after 
> an update. The update endpoint returns success, but a subsequent 
> GET sometimes shows the old data.
>
> Think through possible causes step by step. Our stack:
> - Express API with PostgreSQL
> - Redis cache with 5-minute TTL
> - Read replicas for GET requests
> - Nginx reverse proxy with caching headers

The AI reasons through:

Possible causes (ordered by likelihood):
1. Redis cache not invalidated on update → most common cause
2. Read replica lag → PostgreSQL replication delay
3. Nginx caching the GET response → check Cache-Control headers
4. Client-side caching → browser/CDN caching the response
5. Race condition → GET arrives before UPDATE commits

Diagnostic steps:
1. Check: Does the problem resolve after 5 minutes? → Redis TTL
2. Check: Bypass cache and query primary directly → replica lag
3. Check: curl with no-cache headers → Nginx/client cache
4. Check: Does it only happen under load? → race condition

Most likely: #1 (Redis cache). The update endpoint probably 
doesn't invalidate the cached user object.

Example: Intermittent test failure

> This test passes 90% of the time but occasionally fails with 
> "expected 3 items but received 2". Think through why a test 
> might be flaky in this way:
>
> The test creates 3 records, then queries for all records 
> and asserts the count.

Deepening the investigation:

> You identified that another test might be cleaning up records.
> Think through: how do I confirm this? What would I look for 
> in the test output? How do I fix it without slowing down the 
> test suite?

Pattern: "What would a senior engineer check?"

> A senior engineer is debugging this memory leak. 
> Walk through their thought process:
> - What metrics would they look at first?
> - What tools would they use?
> - What would they test and in what order?
> - When would they escalate vs keep investigating?
WARNING
Watch Out
CoT for debugging works best when you provide concrete symptoms and stack details. Vague descriptions like 'something is broken' give the AI nothing to reason about. Include error messages, logs, and what you've already tried.

4 When CoT Helps vs When It Doesn't

Chain-of-thought is powerful, but it's not always the right tool. Here's a guide:

CoT is most valuable for:

Situation Why CoT Helps
Multi-factor decisions Forces consideration of all factors
Debugging complex issues Systematic investigation vs random guessing
Architecture choices Tradeoff analysis needs explicit reasoning
Code review Explaining why something is problematic
Input validation Tracing every place external data enters needs step-by-step reasoning
Performance optimization Need to identify bottlenecks before optimising

CoT is unnecessary for:

Situation Why Skip CoT
Simple code generation "Write a function to sort an array" — just do it
Formatting/styling "Convert these to camelCase" — no reasoning needed
Boilerplate generation "Create a React component" — pattern is straightforward
Looking up facts "What's the Express middleware signature?" — just answer
Small, clear tasks "Add a null check before line 15" — obvious intent

CoT can actually hurt when:

  • The task is simple and CoT adds unnecessary verbosity
  • You need quick iteration and the reasoning slows output
  • The problem is well-defined with one obvious solution
  • You're generating large amounts of code and don't need justification for each line

Choosing the right level of reasoning:

No CoT:     "Write a function that validates email addresses"
Light CoT:  "Write an email validator and briefly explain your regex choice"
Medium CoT: "Design an email validation strategy — consider edge cases before implementing"
Heavy CoT:  "Think through email validation thoroughly: what standards exist, 
             what tradeoffs between strict/permissive, what are the edge cases, 
             then implement the best approach for a user registration form"

Rule of thumb: If you'd want a colleague to explain their reasoning before implementing, use CoT. If you'd just say "go do it", skip CoT.

TIP
Tip
Start without CoT. If the AI gives a shallow or wrong answer, re-prompt with CoT. This avoids over-engineering simple requests while ensuring complex ones get proper thought.

5 Structured Thinking Templates

Here are reusable CoT templates for common development scenarios. Adapt these to your specific needs.

Template: Technology Evaluation

> Evaluate [technology] for our use case. Think through:
> 1. What problem does it solve that we currently have?
> 2. What's the learning curve for our team?
> 3. What's the operational burden (hosting, monitoring, upgrades)?
> 4. What are we giving up by adopting it?
> 5. What's the exit strategy if it doesn't work out?
> 6. Based on this analysis, should we adopt it? For what scope?

Template: Code Design Decision

> Before implementing, think through the design:
> 1. What are the inputs and outputs of this system?
> 2. What are the key operations and their frequency?
> 3. What are the failure modes and how should each be handled?
> 4. What will change most often in the future?
> 5. Design the interfaces first, then describe the implementation.

Template: Bug Root Cause Analysis

> Help me find the root cause. Think step by step:
> 1. What is the exact symptom? (What happens vs what should happen)
> 2. When did it start? (What changed recently?)
> 3. What are all possible causes? (List at least 5)
> 4. Rank causes by likelihood given the evidence
> 5. For the top cause, what confirms or denies it?
> 6. Suggest a fix and explain why it addresses the root cause

Template: Performance Investigation

> This operation is slow ([current time] vs [target time]). 
> Think through the performance analysis:
> 1. What are all the operations happening in this path?
> 2. Which operations are likely bottlenecks? (I/O, computation, network)
> 3. What would a flame graph likely show?
> 4. What's the quickest win vs the most impactful fix?
> 5. Are there caching opportunities? At what layer?

Template: Input Handling Review

> Review how this code handles external input. Think through:
> 1. Where does data from outside the code enter?
> 2. What unexpected input could it receive? (empty, huge, wrong type, wrong encoding)
> 3. Is every query parameterised, and every value validated before use?
> 4. What breaks, and how far does the damage spread, if one input is bad?
> 5. Where would a second check catch what the first one misses?

Combining templates in practice:

You can chain templates for thorough analysis:

> First, use the bug root cause template to identify the issue.
> Then, use the code design template to design a proper fix.
> Finally, think through whether the fix introduces any new input-handling problems.
NOTE
Key Insight
These templates work because they force both you and the AI to think in a structured way. The numbered steps create a logical flow that's harder to skip over or shortcut.

Questions & Answers

Q: Does "think step by step" still work, or is it outdated?
It still works but it's quite generic. More specific prompts like "analyse the tradeoffs between X and Y before deciding" produce better targeted reasoning. The generic phrase is a fine fallback when you don't know what specific thinking to request.
Q: Does CoT use more tokens / cost more?
Yes — the AI generates more output when reasoning step by step, which uses more tokens. For simple tasks, this is wasteful. For complex decisions, the cost is negligible compared to the value of getting the right answer.
Q: Should I put CoT instructions in my system prompt?
Selectively. You might add "when I ask about architecture decisions, always think through tradeoffs before recommending" to your rules file. But don't add "always think step by step" globally — it will make simple tasks verbose.
Q: Can I tell the AI to think internally without showing me?
Some tools support this (Claude's "extended thinking" feature reasons internally before responding). When available, this gives you the benefits of CoT without the verbose output. Check your tool's documentation for this capability.

Key Takeaways

  1. CoT improves complex decisions — step-by-step reasoning catches flaws that direct answers miss
  2. Be specific about what to think through — "analyse tradeoffs of X vs Y" beats "think step by step"
  3. Use CoT for debugging — systematic investigation beats random guessing
  4. Skip CoT for simple tasks — it adds verbosity without value for straightforward requests
  5. Build templates — reusable CoT structures ensure thorough analysis every time
  6. Verify the reasoning — CoT lets you check the logic, not just the conclusion

Next Steps: In Lesson 8 — Multi-Turn Conversations, you'll learn how to build effective context across multiple messages for larger, more complex tasks.