"Build better evals" is the most repeated advice in AI engineering. The hard part is doing it when the output is a slide deck. In 45 minutes you'll wire up a Managed Agent that generates decks, score it against SlidesBench, and iterate the prompt based on what fails. You'll leave knowing how to turn "this looks bad" into a number you can move, with a working eval loop to prove it.

Keynotes, demos, and conversations with the teams behind Claude. Recorded at Code w/ Claude 2026 San Francisco and ready to replay.