AI Evals: Everything You Need to Know
A long, free FAQ from two practitioners who teach evals to thousands of engineers, covering error analysis, when to use an LLM as a judge, annotation, RAG and agent evaluation, and what to run in CI versus production. Its central argument, that reading your own failures matters more than any metric, is the one most teams learn too late.
Guide
AI & Machine Learning
Hamel Husain and Shreya Shankar
Intermediate
Free
A long, free FAQ from two practitioners who teach evals to thousands of engineers, covering error analysis, when to use an LLM as a judge, annotation, RAG and agent evaluation, and what to run in CI versus production. Its central argument, that reading your own failures matters more than any metric, is the one most teams learn too late.
Ready to start?
Opens on Hamel Husain and Shreya Shankar in a new tab.
Stuck on something specific?
Writing only gets you so far. If you want an answer to your situation rather than the general case, book a session and we will work through it together. Sessions are free for approved Sefism members, and a few slots open each week.
Follow along
New writing, resources and project ideas land here first.