A live walkthrough of building an eval from real product data.
Why AI evals are the hottest new skill for product builders
Lenny's Podcast (Hamel Husain & Shreya Shankar) Sep 2025
Watch on YouTube youtube.com →You test agents with evals: a set of real tasks plus grading logic, run repeatedly because the same input can produce different outputs. Start embarrassingly simple: read 50 transcripts of your agent working, label every failure, turn the recurring ones into automated checks (code checks where possible, an LLM judge validated against your own labels where judgment is needed). Trust is earned per task: track pass rates over many runs, not one good demo.
18 resources.
A live walkthrough of building an eval from real product data.
Lenny's Podcast (Hamel Husain & Shreya Shankar) Sep 2025
Watch on YouTube youtube.com →The same masterclass in podcast form for the commute.
Lenny's Podcast Sep 2025
Listen on Spotify open.spotify.com →The honest failure modes of eval culture, so you avoid cargo-culting it.
Vanishing Gradients (Hamel Husain) 2025
Open vanishinggradients.fireside.fm →A second full session using messy real data instead of toy examples.
Product Growth (Hamel & Shreya) 2025
Open news.aakashg.com →The clearest official explanation of what an agent eval actually is.
Anthropic 2025
Open anthropic.com →Every question you will have about evals, answered by the field's top teacher.
Hamel Husain 2025
Open hamel.dev →The philosophy behind the top-grossing evals course, free to read.
Hamel Husain 2025
Open hamelhusain.substack.com →The course that trained 2,000+ people, including teams at OpenAI and Anthropic.
Hamel Husain & Shreya Shankar 2025
Open maven.com →The 60-second version of why evals became the must-have skill.
Lenny Rachitsky Sep 2025
Open x.com →Written notes of the full workflow: error analysis to automated judges.
Aakash Gupta 2025
Open aakashg.com →Explains pass@k and why you must run agent tests many times.
Cameron Wolfe 2025
Open cameronrwolfe.substack.com →Field lessons from teams that moved agents from demo to production.
InfoQ 2025
Open infoq.com →A single-sitting starter recipe for your first agent eval.
Zen van Riel 2025
Open zenvanriel.com →Covers the metrics that matter: task success, step count, faithfulness.
Turing College 2025
Open turingcollege.com →Connects evals to shipping cadence rather than treating them as QA theatre.
Lenny's Newsletter 2025
Open lennysnewsletter.com →A worked example you can copy line by line for your own product.
Arize AI 2025
Open arize.com →Surveys the tooling landscape so you do not build eval infrastructure from scratch.
Dave Davies 2025
Open medium.com →A checklist to run before you let an agent near customers.
Maxim AI 2025
Open getmaxim.ai →The same ground, over in Build the product, our Starting Up track.