Demystifying evals for AI agents

Anthropic··Submitted by Mads Kristian Nylund
kpiagentsanthropicevaluation

Anthropic lays out its field-tested framework for evaluating AI agents: the vocabulary (tasks, trials, graders, transcripts, harnesses), the three grader types (code-based, model-based, human), capability vs. regression evals, and two key non-determinism metrics, pass@k (odds of at least one success in k attempts) and pass^k (odds all k attempts succeed), plus an 8-step roadmap for building an eval suite from scratch.

Read Article

More from Anthropic

Related Articles