
Agent observability is crucial for evaluating AI agents due to their non-deterministic behavior, which traditional software observability cannot capture. Evaluations differ from software testing by focusing on reasoning, context, and state evolution across multi-turn interactions. Traces provide data for assessing agent behavior at various granularities, such as single steps or full conversations, and are essential for diagnosing issues in real-world scenarios. The evaluation process requires selecting appropriate granularity levels, with full-turn evaluations relying on detailed trace data and single-step evaluations using run captures. Ensuring trace data captures all necessary context is critical for accurate assessment of agent performance.
Managed Deep Agents gives developers a managed way to build, run, and deploy Deep Agents with built-in runtime, streaming, sandboxes, evals, memory, and auth.

LangSmith Bring Your Own Cloud is now generally available on AWS, giving Enterprise teams managed observability, evaluation, and deployment inside their own VPC.

Learn what AI agents are, how they work in an LLM loop, and where workflows fit so you can build reliable, production-ready autonomous systems.
