
LangSmith's pairwise evaluation enables users to compare two LLM-generated responses using a custom evaluator, offering a clearer assessment of which output is better based on predefined criteria. This feature is particularly useful for tasks without a single correct answer, such as content generation, and allows for direct comparison of two results. In an example, pairwise evaluation revealed preferences among four LLMs in generating engaging tweets, demonstrating its effectiveness in highlighting differences between models.
Managed Deep Agents gives developers a managed way to build, run, and deploy Deep Agents with built-in runtime, streaming, sandboxes, evals, memory, and auth.

LangSmith Bring Your Own Cloud is now generally available on AWS, giving Enterprise teams managed observability, evaluation, and deployment inside their own VPC.

Learn what AI agents are, how they work in an LLM loop, and where workflows fit so you can build reliable, production-ready autonomous systems.
