On demand

Evals for AI Agents: How Product Builders Get the Most Out of Every New Model

Join Anthropic's Applied AI team for a session on evaluating AI agents.

Most teams shipping AI agents can't tell whether a new model actually improves their product.The hard part is that the products people are shipping now aren't single prompts. They're agents that call tools, retrieve context, and take several steps before a customer sees output, with a different way to fail at each step. We'll show how the teams we work with evaluate agents end to end, decide quickly whether a new model is worth switching to, and stay confident their product performs at the level customers expect.

You'll see real examples from startups building on Claude and leave with something you can put into practice the same week.

Featured speakers

  • Preston Tuggle

    Applied AI @ Anthropic

  • Jimmy Chan

    Applied AI @ Anthropic

What you’ll learn

  • Why single-turn evals miss most of what goes wrong in an AI agent, and what to measure instead

  • How to build a first eval set from real production failures

  • How to decide quickly whether a new model release is worth adopting for your product

Transform how your organization operates with Claude