
Evaluating voice agents involves three key dimensions: execution, outcome, and experience. Execution focuses on whether the agent followed instructions, including tool calls and policies. Outcome assesses if the interaction achieved its goal, even if instructions were followed. Experience measures caller satisfaction through factors like responsiveness, naturalness, and clarity. A continuous evaluation loop, such as in LangSmith, allows developers to compare prompts, models, and workflows against representative conversations, ensuring consistent performance across platforms.
Managed Deep Agents gives developers a managed way to build, run, and deploy Deep Agents with built-in runtime, streaming, sandboxes, evals, memory, and auth.

LangSmith Bring Your Own Cloud is now generally available on AWS, giving Enterprise teams managed observability, evaluation, and deployment inside their own VPC.

Learn what AI agents are, how they work in an LLM loop, and where workflows fit so you can build reliable, production-ready autonomous systems.
