Demystifying evals for AI agents

Anthropic describes the engineering behind their multi-agent Research feature, where a planning agent decomposes complex queries and spawns parallel search agents. The post covers architectural principles, prompting strategies, and evaluation methods for reliable multi-agent systems.
