Agentic Engineering Handbook
The definitive OpenAI, Anthropic, Google, MCP, Harness, Evals, and Production Agent Systems learning roadmap.
If this repository helps you, consider giving it a ⭐
Why This Repository?
The AI industry has entered the Agentic Era. Building production-grade AI systems now requires mastering agents, tool use, MCP, memory, long-running workflows, coding agents, agent harnesses, evals, and safety — but the knowledge is scattered across OpenAI blogs, Anthropic engineering posts, SDK docs, cookbooks, and research papers.
This repository consolidates 175 curated resources into one structured learning roadmap.
The goal: Become a world-class Agentic Engineer.
How To Use This Handbook
Pick the path that matches your starting point:
- New to agents: follow the Learning Roadmap from Phase 0 to Phase 6. Treat each
Read First,Then Read, andBuild Exerciseas a checklist. - Already building LLM apps: start at Phase 2 or Phase 3, then fill gaps in agent loop, tool calling, evals, and production engineering.
- Trying to build projects: use the phase-level
Build Exerciseprompts, then branch into Applied Practice Tracks for coding agents, security, code review, or SRE. - Looking for references: jump to the Full Reading Table. Read
P0first, useP1for implementation detail, and keepP2as optional background.
Learning Roadmap
Phase 0 — Agent Loop From Scratch
If you treat Claude Code as a coding CLI, many capabilities can feel like magic: it reads files, runs commands, edits code, delegates work, and stays oriented during complex tasks.
From an engineering perspective, the core is much simpler:
model + tools + one loop.
Understanding that loop makes the rest of the system easier to reason about:
- When the agent should plan first, and when it should act immediately
- Why an explicit todo list reduces drift in longer tasks
- Why subagents improve exploration while protecting the main context
- How skills, MCP, and hooks each add capability around the same core loop
These pages are based on the upstream English Markdown tutorials from shareAI-lab/mini-claude-code, with added Study Notes and inline source code for this handbook.
Supporting files are included in the same folder: requirements.txt, .env.example, v0_bash_agent_mini.py, and skills/.
Phase 1 — Agent Foundations
Build shared vocabulary for workflow vs agent, tool loop, handoff, guardrails.
Key Mental Models
Should I build an agent? (4-question checklist from Barry Zhang's talk)
| Question | If No → Workflow | If Yes → Agent |
|---|---|---|
| Is the task complex enough? | Decision tree is fully mappable | Ambiguous problem space |
| Is the task valuable enough? | <$0.10 per run | >$1 per run, cost doesn't matter |
| Are all core capabilities doable? | Weak links break the chain | Model handles every step well |
| Is error cost low & detectable? | High cost + hard to detect → human-in-the-loop | Errors caught by tests/CI |
Think like the agent. Most failures come from designing with a human perspective. Put yourself inside the agent's context window: you only see ~10K–20K tokens (system prompt + tool descriptions + recent observations). Ask: does the agent have enough information to act correctly at each step?
→ Source: How We Build Effective Agents
Read First
| # | Title | Vendor |
|---|---|---|
| 1 | System Prompts | Anthropic |
| 2 | Prompt guidance | OpenAI |
| 3 | Function Calling | OpenAI |
| 4 | Tool use overview | Anthropic |
| 5 | Function calling - Gemini API | |
| 6 | Building effective agents | Anthropic |
| 7 | New tools for building agents | OpenAI |
| 8 | Agents SDK overview | OpenAI |
Then Read
| Title | Vendor |
|---|---|
| How We Build Effective Agents: Barry Zhang, Anthropic | Anthropic |
| Phistory — Claude Code & Codex CLI System Prompt Diff History | Community |
| Coding Agents 101: The Art of Actually Getting Things Done | Cognition |
| OpenAI Agents SDK examples | OpenAI |
| Structured Outputs for Multi-Agent Systems | OpenAI |
Build Exercise
Build a customer service/ticket triage agent: router → specialist → evaluator, with all outputs constrained by structured schemas.
Phase 2 — MCP & Tool Ecosystem
Understand MCP server/client, remote vs local, tool loading, approval, connector boundaries.
Read First
| # | Title | Vendor |
|---|---|---|
| 1 | Introducing the Model Context Protocol | Anthropic |
| 2 | MCP and Connectors | OpenAI |
| 3 | Building MCP servers for ChatGPT Apps and API integrations | OpenAI |
Then Read
| Title | Vendor |
|---|---|
| Code execution with MCP: Building more efficient agents | Anthropic |
| Writing effective tools for AI agents - with AI agents | Anthropic |
| Model Context Protocol - Codex | OpenAI |
| Build a Remote MCP server | Cloudflare |
| Introducing the MCP Registry | MCP |
| OpenAI Docs MCP | OpenAI |
| Build your ChatGPT UI | OpenAI |
Build Exercise
Build a read-only repo/docs MCP server, then create an eval to verify the agent correctly cites documentation.
Phase 3 — Context, Memory & Skills
Learn to control context window, short/long-term memory, skills/plugins, CLAUDE.md/AGENTS.md.
Read First
| # | Title | Vendor |
|---|---|---|
| 1 | Agent Skills Specification | Agent Skills |
| 2 | Effective context engineering for AI agents | Anthropic |
| 3 | How the Open Knowledge Format can improve data sharing | Google Cloud |
| 4 | How Long Contexts Fail | Drew Breunig |
| 5 | Context Rot | Chroma |
| 6 | Progressive disclosure | Claude-Mem |
| 7 | Equipping agents for the real world with Agent Skills | Anthropic |
| 8 | Agent Skills | Anthropic |
| 9 | Skills | OpenAI |
| 10 | Building Reliable Agents with Memory and Compaction | OpenAI |
Then Read
| Title | Vendor |
|---|---|
| Custom instructions with AGENTS.md - Codex | OpenAI |
| Best practices for Claude Code | Anthropic |
| Agent Skills - Codex | OpenAI |
| Skills in OpenAI API | OpenAI |
Build Exercise
Implement the same task as a Skill/Plugin, then measure accuracy and token cost across three variants: no skill, long prompt, and skill-based.
Phase 4 — Harness & Long-Running Agents
Master agent runtime: event stream, thread, tool execution, state, sandbox, approval, recovery.
Read First
| # | Title | Vendor |
|---|---|---|
| 1 | Unrolling the Codex agent loop | OpenAI |
| 2 | Unlocking the Codex harness: how we built the App Server | OpenAI |
| 3 | Agent Harness Engineering: A Survey | Academic |
| 4 | Effective harnesses for long-running agents | Anthropic |
| 5 | Deep Agents | LangChain |
Then Read
| Title | Vendor |
|---|---|
| Deep research | OpenAI |
| Open Deep Research | LangChain |
| The next evolution of the Agents SDK | OpenAI |
| Using PLANS.md for multi-hour problem solving | OpenAI |
| Build long-running AI agents that pause, resume, and never lose context with ADK | |
| Harness design for long-running application development | Anthropic |
| Scaling Managed Agents: Decoupling the brain from the hands | Anthropic |
Build Exercise
Build a mini coding harness: plan file, shell tool, apply patch, test gate, event log, and resume capability.
Phase 5 — Coding & Workspace Agents
Compare Codex vs Claude Code product/SDK forms; learn multi-agent, IDE, workspace collaboration.
Read First
| # | Title | Vendor |
|---|---|---|
| 1 | AGENTS.md | Agentic AI Foundation |
| 2 | Introducing Codex | OpenAI |
| 3 | Best practices for Claude Code | Anthropic |
| 4 | How Claude Code works in large codebases | Anthropic |
| 5 | Enabling Claude Code to work more autonomously | Anthropic |
Then Read
| Title | Vendor |
|---|---|
| Introducing the Codex app | OpenAI |
| Introducing workspace agents in ChatGPT | OpenAI |
| Apple's Xcode now supports Claude Agent SDK | Anthropic |
| Building Consistent Workflows with Codex CLI & Agents SDK | OpenAI |
| Best practices for Claude Code | Anthropic |
| The spec is dead, long live the spec! | Ravi on Product |
| How Anthropic teams use Claude Code | Anthropic |
| Multi-stack Web App Builds | Community |
Build Exercise
Run both OpenAI/Codex and Claude Code style workflows on the same repo: issue → plan → patch → tests → PR summary.
Phase 6 — Evals, Safety & Production
Build pre/post-launch eval loop, trace loop, safety boundaries, permissions, regression monitoring.
Read First
| # | Title | Vendor |
|---|---|---|
| 1 | Demystifying evals for AI agents | Anthropic |
| 2 | The six generations of AI agents and how to eval them | Braintrust |
| 3 | Agent observability powers agent evaluation | LangChain |
| 4 | Agent Evaluation Readiness Checklist | LangChain |
| 5 | Build an Agent Improvement Loop with Traces, Evals, and Codex | OpenAI |
| 6 | Macro Evals for Agentic Systems | OpenAI |
| 7 | Testing Agent Skills Systematically with Evals | OpenAI |
Then Read
| Title | Vendor |
|---|---|
| How we build evals for Deep Agents | LangChain |
| Deep Research Bench | FutureSearch |
| How to Evaluate Tool-Calling Agents | Arize |
| AI agent evaluation: How to test, debug, and improve agents in production | Arize |
| A Survey on Agent-as-a-Judge | Academic |
| Running Codex safely at OpenAI | OpenAI |
| How we contain Claude across products | Anthropic |
| Evals API Use-case - MCP Evaluation | OpenAI |
| Measuring AI agent autonomy in practice | Anthropic |
Build Exercise
Build a smoke/macro eval suite for your agent: task success rate, tool misuse, prompt injection resistance, latency, cost, and human approval count.
Applied Practice Tracks
Use these tracks after the core roadmap when you want to practice agentic engineering in real engineering workflows.
| Track | Start Here | Why It Matters |
|---|---|---|
| Agentic coding workflow | Coding Agents 101, How Claude Code works in large codebases, How Anthropic teams use Claude Code | Turns agent theory into day-to-day engineering habits: prompting, checkpoints, verification, parallel work, and team rollout. |
| Spec-driven building | The spec is dead, long live the spec!, Multi-stack Web App Builds | Treats specs, prompts, and assignments as executable source material for agents. |
| Context failure modes | How Long Contexts Fail, Context Rot, Progressive disclosure | Helps diagnose context poisoning, distraction, confusion, context degradation, and retrieval overload. |
| Evals and observability | Demystifying evals for AI agents, Agent observability powers agent evaluation, Agent Evaluation Readiness Checklist | Builds the feedback loop for traces, datasets, graders, offline/online evals, and regression gates. |
| Deep research agents | Deep research, Open Deep Research, Alibaba-NLP/DeepResearch | Practices long-running research agents: planning, search, MCP, citations, report synthesis, and benchmark-driven improvement. |
| MCP operations | Build a Remote MCP server, Introducing the MCP Registry | Shows how MCP moves from local prototypes to authenticated, discoverable, production-grade tool ecosystems. |
| Agent security | OWASP Top Ten, SAST vs. DAST vs. RASP, Copilot Remote Code Execution via Prompt Injection | Grounds agent security in classic AppSec plus new prompt-injection and tool-permission failure modes. |
| Code review systems | How to Review Code Effectively, AI-Assisted Assessment of Coding Practices in Modern Code Review, AI Code Review Implementation Best Practices | Connects human review quality with AI-assisted review, automated comments, and review policy design. |
| Production and SRE agents | ML and LLM system design, Introduction to Site Reliability Engineering, Observability Basics You Should Know | Extends agents beyond coding into incidents, observability, root-cause analysis, on-call, and production operations. |
Full Reading Table
Priority guide: P0 = must-read (architectural/conceptual), P1 = highly useful (implementation detail), P2 = optional context (background/releases).
| Priority | Title | Vendor | Topic | Key Idea | Date |
|---|---|---|---|---|---|
| P0 | OpenAI for Developers in 2025 | OpenAI | Agents; MCP; Platform | Annual overview: systematic walkthrough of Responses API, Agents SDK, AgentKit, Codex, MCP, Apps SDK, and AGENTS.md. | 2025-12-30 |
| P0 | New tools for building agents | OpenAI | Agents; Responses API; Tools | Key starting point for OpenAI's agent platform: Responses API, built-in web/file/computer tools, Agents SDK, tracing/observability. | 2025-03-11 |
| P0 | Introducing AgentKit | OpenAI | Agents; Evals; AgentKit | AgentKit, expanded evals, agent RFT: the official agent toolchain from prototype to production. | 2025-10-06 |
| P0 | Prompt guidance | OpenAI | Prompting; Models; Agent UX | Official model-specific prompting guidance for outcome-first prompts, reasoning effort, preambles, and validation rules in tool-heavy workflows. | Current docs |
| P0 | System Prompts | Anthropic | System prompts; Claude; Behavior | Claude web/mobile system prompt release notes; useful for studying production prompting patterns and behavioral scaffolding. | Current docs |
| P0 | Agents SDK overview | OpenAI | Agents; SDK | Official SDK entry point: concepts and boundaries of agent, tool, handoff, guardrail, and tracing. | Current docs |
| P0 | Introducing the Model Context Protocol | Anthropic | MCP; Standards | The origin article for MCP: an open standard connecting AI assistants to data, tools, and systems. | 2024-11-25 |
| P0 | Building effective agents | Anthropic | Agents; Patterns; Frameworks | Essential agent primer: workflow vs agent, prompt/tool/retrieval, orchestrator-worker, evaluator-optimizer patterns. | 2024-12-19 |
| P0 | Coding Agents 101: The Art of Actually Getting Things Done | Cognition | Coding agents; Workflows; Practice | Product-agnostic guide to prompting, delegation, verification, environment setup, security, and cost management for coding agents. | 2025-06 |
| P0 | AGENTS.md | Agentic AI Foundation | Coding agents; Repo instructions; Standards | Open Markdown convention for giving coding agents setup, test, style, safety, and workflow instructions across repos and tools. | Current docs |
| P0 | New tools and features in the Responses API | OpenAI | MCP; Responses API; Tools | Responses API extended to remote MCP servers, image/code/file tools; see how OpenAI integrates MCP into its runtime. | 2025-05-21 |
| P0 | MCP and Connectors | OpenAI | MCP; Connectors; Responses API | Official guide to connecting remote MCP servers and connectors; includes approvals and security considerations. | Current docs |
| P0 | Building MCP servers for ChatGPT Apps and API integrations | OpenAI | MCP; ChatGPT Apps; API | Official guide to writing MCP servers: supply tools/knowledge to ChatGPT Apps, deep research, and API integrations. | Current docs |
| P0 | Deep research | OpenAI | Deep research; MCP; API | Official guide to deep research models, including web search, file search, remote MCP servers, code interpreter, and security risks. | Current docs |
| P0 | Building a Deep Research MCP Server | OpenAI | MCP; Deep research | Minimal implementation of a search/fetch MCP server for Deep Research. | 2025-06-25 |
| P0 | Model Context Protocol - Codex | OpenAI | MCP; Codex | How Codex CLI/IDE connects to MCP servers, adding Figma, browser, docs, and internal tool context to agents. | Current docs |
| P0 | Introducing Codex | OpenAI | Agents; Coding; Sandbox | Cloud-based software engineering agent: parallel tasks, repo sandbox, running tests/linters/type checkers, producing auditable evidence. | 2025-05-16 |
| P0 | Agent Harness Engineering: A Survey | Academic | Harness; Taxonomy; Agent architecture | Survey that frames harness engineering as its own system layer and introduces the ETCLOVG taxonomy: Execution, Tooling, Context, Lifecycle, Observability, Verification, and Governance. | 2026 |
| P0 | Unrolling the Codex agent loop | OpenAI | Harness; Agent loop; Codex | How Codex CLI chains prompt, tool schema, MCP tools, Responses API, and context management into an agent loop. | 2026-01-23 |
| P0 | Unlocking the Codex harness: how we built the App Server | OpenAI | Harness; Codex App Server; JSON-RPC | Core harness article: Codex core, App Server, JSON-RPC, streaming progress, approval, diff, and thread management. | 2026-02-04 |
| P0 | From model to agent: Equipping the Responses API with a computer environment | OpenAI | Harness; Responses API; Sandbox | Responses API + shell tool + hosted containers form the agent runtime; essential for understanding the model-to-agent execution environment. | 2026-03-10 |
| P0 | Harness engineering: leveraging Codex in an agent-first world | OpenAI | Harness; Agent-first engineering | Design product code, tests, CI, docs, and observability to be agent-readable/executable; learn agent-first repo organization. | 2026-02-11 |
| P0 | The next evolution of the Agents SDK | OpenAI | Harness; Agents SDK; MCP; Skills | Agents SDK harness becomes more complete: memory, sandbox orchestration, Codex-like filesystem tools, MCP, skills, AGENTS.md. | 2026-04-15 |
| P0 | Building Consistent Workflows with Codex CLI & Agents SDK | OpenAI | MCP; Codex; Agents SDK | Codex CLI as an MCP server integrated with Agents SDK; real multi-agent dev workflow. | 2025-10-01 |
| P0 | Building Reliable Agents with Memory and Compaction | OpenAI | Memory; Compaction; Reliability | Memory and compaction design for long-context/multi-turn agents. | 2026-05-01 |
| P0 | Build an Agent Improvement Loop with Traces, Evals, and Codex | OpenAI | Evals; Traces; Self-improvement | Connect traces, evals, and Codex fixes into an agent improvement loop. | 2026-05-12 |
| P0 | Eval Driven System Design - From Prototype to Production | OpenAI | Evals; Production | Use evals as the driving force for system design; ideal for moving agents from demo to production. | 2025-06-02 |
| P0 | Testing Agent Skills Systematically with Evals | OpenAI | Evals; Skills; Agents | Systematically test agent skills with evals; establish quality gates before skill release. | 2026-01-22 |
| P0 | Evals API Use-case - MCP Evaluation | OpenAI | MCP; Evals | Evaluate QA/retrieval capabilities with MCP tools; ideal for building an MCP regression suite. | 2025-06-09 |
| P0 | The six generations of AI agents and how to eval them | Braintrust | Evals; Agent architecture; Harness | Maps six generations of agent architecture to the eval strategy each generation requires, from prompts to AI harnesses. | 2026-05-21 |
| P0 | Agent observability powers agent evaluation | LangChain | Evals; Observability; Traces | Explains why traces are the source of truth for agent behavior and how observability feeds evaluation. | 2026-01-27 |
| P0 | Running Codex safely at OpenAI | OpenAI | Safety; Sandbox; Codex | How OpenAI runs Codex internally: sandbox, approvals, network policy, agent-native telemetry. | 2026-05-20 |
| P0 | Building Governed AI Agents - A Practical Guide to Agentic Scaffolding | OpenAI | Governance; Guardrails; Agents | Governed agent scaffolding: permissions, guardrails, auditing, and organizational policies. | 2026-02-23 |
| P0 | Macro Evals for Agentic Systems | OpenAI | Evals; Agentic systems | Evaluate agents at the end-to-end/macro level, not just individual step outputs. | 2026-05-19 |
| P0 | Best practices for Claude Code | Anthropic | Coding agents; Claude Code | Claude Code methodology: verification loop, explore-plan-code, CLAUDE.md, permissions, MCP, subagents, context management. | 2025-04-18 |
| P0 | Best practices for Claude Code | Anthropic | Claude Code; Coding agents; Workflow | Official Claude Code docs for planning, CLAUDE.md, verification, tool use, and team workflows. | Current docs |
| P0 | How Claude Code works in large codebases | Anthropic | Claude Code; Large codebases; Enterprise | Patterns for large-codebase Claude Code adoption: layered CLAUDE.md, hooks, skills, plugins, MCP, LSP, subagents, and rollout ownership. | 2026-05-14 |
| P0 | How we built our multi-agent research system | Anthropic | Agents; Multi-agent; Research | Claude Research multi-agent architecture: planner + parallel research agents + synthesis; production multi-agent experience. | 2025-06-13 |
| P0 | Writing effective tools for AI agents - with AI agents | Anthropic | Tools; MCP; Evals | Tool quality determines agent quality: tool descriptions, context budget, eval, and letting Claude optimize its own tools. | 2025-09-11 |
| P0 | Effective context engineering for AI agents | Anthropic | Context; Agents | Context is the agent's core resource: selection, compression, isolation, persistence, and context pollution control. | 2025-09-29 |
| P0 | How Long Contexts Fail | Drew Breunig | Context; Long context; Agents | Taxonomy of context poisoning, distraction, confusion, and clash; practical fixes for overloaded agent contexts. | 2025-06-22 |
| P0 | Context Rot | Chroma | Context; Long-context evals | Research on how LLM performance degrades as input grows, especially with distractors and similar-but-wrong context. | 2025-07-16 |
| P0 | Enabling Claude Code to work more autonomously | Anthropic | Claude Code; Agent SDK; Subagents | Claude Agent SDK, subagents, hooks, background tasks, checkpoints, and other autonomous coding agent capabilities. | 2025-09-29 |
| P0 | Equipping agents for the real world with Agent Skills | Anthropic | Skills; Agents | Agent Skills as modular capability packages: instructions, resources, scripts — reducing context burden and improving reliability. | 2025-10-16 |
| P0 | Agent Skills | Anthropic | Skills; Claude; Progressive disclosure | Official Claude Agent Skills docs: modular instructions, metadata, scripts, resources, and on-demand loading across Claude products. | Current docs |
| P0 | Skills | OpenAI | Skills; API; Shell environments | Official OpenAI API guide for uploading, managing, and attaching reusable Skills to hosted and local shell environments. | Current docs |
| P0 | Agent Skills Specification | Agent Skills | Skills; Specification; Progressive disclosure | Complete skill package format: SKILL.md frontmatter, optional scripts/references/assets, file references, and validation. | Current docs |
| P0 | Code execution with MCP: Building more efficient agents | Anthropic | MCP; Code execution; Context | Key article on MCP scale challenges: reduce token overhead with code execution/on-demand tools; learn progressive disclosure. | 2025-11-04 |
| P0 | Introducing advanced tool use on Claude Developer Platform | Anthropic | Tools; MCP; Advanced tool use | Tool search, deferred loading, programmatic tool calling; solving context pollution from large numbers of MCP tools. | 2025-11-24 |
| P0 | Effective harnesses for long-running agents | Anthropic | Harness; Long-running agents | Essential harness reading: working across multiple context windows, task logging, external state, agent self-recovery. | 2025-11-26 |
| P0 | Deep Agents | LangChain | Harness; Long-running agents; Deep research | Opinionated open-source agent harness for planning, context management, subagents, filesystem, memory, and human-in-the-loop workflows. | Current repo |
| P0 | Demystifying evals for AI agents | Anthropic | Evals; Agents | Agent evals are more complex than static evals: multi-turn, tools, state changes, creative solutions, failure taxonomy. | 2026-01-09 |
| P0 | Measuring AI agent autonomy in practice | Anthropic | Agents; Autonomy; Measurement | Quantify agent autonomy using metrics like task duration and supervision needs; ideal for building autonomy benchmarks. | 2026-02-18 |
| P0 | Harness design for long-running application development | Anthropic | Harness; Application development | Harness design patterns for delegating long-running app development tasks to agents; compare with OpenAI Codex harness. | 2026-03-24 |
| P0 | Scaling Managed Agents: Decoupling the brain from the hands | Anthropic | Managed agents; Harness | Decouple the model brain from execution hands/harness, keeping interfaces stable as the harness evolves. | 2026-04-08 |
| P0 | How we contain Claude across products | Anthropic | Safety; Containment; Agents | Blast radius of powerful agent releases, human-in-the-loop, and containment strategies. | 2026-05-25 |
| P1 | Structured Outputs for Multi-Agent Systems | OpenAI | Agents; Multi-agent; Structured outputs | Use strict schemas to constrain structured messages and handoffs between multiple agents. | 2024-08-06 |
| P1 | Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku | Anthropic | Agents; Computer use | Claude computer use beta starting point: the model uses a computer via screenshots and actions. | 2024-10-22 |
| P1 | Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet | Anthropic | Agents; Coding; Evals | SWE-bench agent scaffolding article: same model performance strongly depends on harness/scaffolding. | 2025-01-06 |
| P1 | Introducing Operator | OpenAI | Agents; Computer use; Safety | Early product form of browser-based agents: model clicks, types, and executes tasks on web pages, emphasizing user confirmation and safety boundaries. | 2025-01-23 |
| P1 | Computer-Using Agent | OpenAI | Agents; Computer use | Understand how CUA combines vision, mouse/keyboard actions, and environment feedback into an agent loop; compare with Claude computer use. | 2025-01-23 |
| P1 | Claude 3.7 Sonnet and Claude Code | Anthropic | Agents; Coding; Claude Code | Early release of Claude Code, marking Claude's entry into the agentic coding tool space. | 2025-02-24 |
| P1 | The think tool: Enabling Claude to stop and think in complex tool use situations | Anthropic | Tools; Reasoning; Agents | Give the model an explicit think tool in complex tool-use chains; learn tool design for policy-heavy/multi-step decisions. | 2025-03-20 |
| P1 | Evaluating Agents with Langfuse | OpenAI | Evals; Agents | Observe and evaluate Agents SDK runs with Langfuse; learn tracing/eval workflows. | 2025-03-31 |
| P1 | Parallel Agents with the OpenAI Agents SDK | OpenAI | Agents; Parallelism; Agents SDK | Parallel agent patterns: decompose tasks, execute in parallel, aggregate results. | 2025-05-01 |
| P1 | Multi-Agent Portfolio Collaboration with OpenAI Agents SDK | OpenAI | Agents; Multi-agent; Portfolio | Multi-agent collaboration business example: research, analysis, combined output. | 2025-05-28 |
| P1 | MCP-Powered Agentic Voice Framework | OpenAI | MCP; Voice; Agents | V |
No comments yet
Be the first to share your take.