04
No Hype AI: Agentic AI
Understand and build AI agents — autonomous systems that plan, act, and learn from their environment.
What separates an agent from a chatbot — the three capabilities (planning, tool use, memory) that make AI autonomous, with real-world examples.
The core agent loops — ReAct, Plan-and-Execute, and Reflexion — with diagrams, pseudocode, and the trade-offs between them.
How agents act on the world: function calling, the Model Context Protocol (MCP), APIs, databases, and robust tool-schema design.
The four kinds of agent memory — short-term, long-term, episodic, and semantic — and how to architect them.
Chain-of-thought, tree-of-thought, and goal decomposition — prompting agents to plan and to re-plan when the environment changes.
Hands-on: build a complete working agent with Claude's tool use — the implementation loop, edge cases, and testing on real tasks.
The safety layers every agent needs — sandboxing, approval loops, resource limits, and output validation — with failure case studies.
Measuring agents beyond accuracy: task completion, efficiency, cost, reliability, and benchmarks like SWE-bench, GAIA, and WebArena.
A pattern catalogue by domain — coding, research, and data agents — their tool sets, prompting strategies, and failure modes.
Open problems and what comes next: long-horizon planning, multi-modal agents, agent collaboration, and emerging standards like MCP.