
IssueBench is an internal benchmark that evaluates the effectiveness of LangSmith Engine by assessing its ability to identify and fix issues in other agents. It focuses on tasks like issue detection, failure categorization, and failure grouping, using synthetic traces with known issues to measure Engine's performance. The benchmark ensures consistency by using a fixed set of issue categories and evaluates the quality of issue cards to support debugging and production behavior. It is designed to test whether Engine has learned underlying failure modes rather than just memorizing surface patterns.
Managed Deep Agents gives developers a managed way to build, run, and deploy Deep Agents with built-in runtime, streaming, sandboxes, evals, memory, and auth.

LangSmith Bring Your Own Cloud is now generally available on AWS, giving Enterprise teams managed observability, evaluation, and deployment inside their own VPC.

Learn what AI agents are, how they work in an LLM loop, and where workflows fit so you can build reliable, production-ready autonomous systems.
