
Evals are systematic methods to assess the quality of LLM outputs based on application-specific criteria. They rely on high-quality data to accurately reflect real-world usage. OpenEvals and AgentEvals offer prebuilt solutions for evaluating conversational quality, writing, and structured outputs. Structured data evaluators ensure output conforms to predefined formats, while LLM-as-a-judge evaluators provide objective scoring without requiring ground truth answers. LangSmith provides tools for tracking and sharing evaluations, aiding in debugging and deploying production-grade LLM applications.
Managed Deep Agents gives developers a managed way to build, run, and deploy Deep Agents with built-in runtime, streaming, sandboxes, evals, memory, and auth.

LangSmith Bring Your Own Cloud is now generally available on AWS, giving Enterprise teams managed observability, evaluation, and deployment inside their own VPC.

Learn what AI agents are, how they work in an LLM loop, and where workflows fit so you can build reliable, production-ready autonomous systems.
