
Test Run Comparisons is a feature in LangChain that enables users to compare multiple test runs side-by-side, providing visual insights into inputs, outputs, and evaluation metrics. It allows for LLM-assisted evaluations and filters to focus on specific datapoints, such as correct or incorrect runs, to identify differences. The feature is designed to help developers debug and understand LLM performance by comparing test runs across iterations. LangSmith, the agent engineering platform, was initially limited in its test comparison capabilities but now offers a user-friendly interface for this purpose, with plans for private beta and broader rollout.
Managed Deep Agents gives developers a managed way to build, run, and deploy Deep Agents with built-in runtime, streaming, sandboxes, evals, memory, and auth.

LangSmith Bring Your Own Cloud is now generally available on AWS, giving Enterprise teams managed observability, evaluation, and deployment inside their own VPC.

Learn what AI agents are, how they work in an LLM loop, and where workflows fit so you can build reliable, production-ready autonomous systems.
