Staging environment

The Hillclimb: How to Measure and Improve Your AI Agents

Hosted by Avery Yip, Lucy Chen, and Colin Matthews

892 students

In this video

What you'll learn

Demystify evals

What evals, benchmarks, and training datasets are, and how they fit together for your product.

Design structured evals

Taxonomize workflows that matter, write verifiable rubric criteria, and build simulations that check what your agent did

Complete the learning loop

Turn failure modes into training data your agent hillclimbs on, so evals become compounding IP instead of a report card.

Why this topic matters

Most teams check their agents on vibes: a few test prompts and sparse top-line metrics. Structured Evals systematically identify where your agent fails, with simulations that mimic real-world scenarios, and turn those failures into the training data that teaches the agent to do the work better. That's how you build AI you can trust. In this session, we'll walk through a real case study: an agent that automates pay disputes.

You'll learn from

Avery Yip

Director of Engineering, Handshake AI

Avery Yip is Director of Engineering at Handshake AI, where he helped build the infrastructure to create data used by frontier AI labs to hillclimb their models, and led the development of agents deployed into Fortune 500 enterprises. Avery joined Handshake AI to build the business from scratch across product and engineering, growing the team to over 50 engineers across five teams and scaling from 0 to over $1B in run rate. More recently, Avery has been building out a new Enterprise business unit focused on helping companies build agents they can trust to deploy into production. Prior to Handshake AI, Avery built the forward-deployed engineering team at Scale AI.

Lucy Chen

Head of Learning and Training, Handshake AI

Lucy Chen leads Learning & Training at Handshake AI, where she teaches domain experts to produce the feedback that makes AI smarter and safer.

She's spent over a decade on how people actually learn. When she's not training models, she's building alternative models of education — homeschooling her triplets.

Colin Matthews

Head of Education, Lenny's Newsletter

I'm a technical PM and builder who has taught more than 40k PMs, designers, and tech professionals the fundamentals of shipping software. I've been featured by Lenny four times, spoken at product conferences and events, and assisted Fortune 500 companies with technical upskilling and AI adoption.


My main goal with any course is to help you build your own mental model of how a technology or system works. Because of how fast the industry is changing, I think its more valuable to understand something from first principles than it is to learn a specific tool or product.

You can check out my courses and free lesson below!

See all products from Colin

Go deeper with a course

Become an AI-Native Builder
Colin Matthews
View syllabus