
The Eval Engineering Skill enables coding agents to generate evaluation tasks (evals) by analyzing agent traces and repository context, allowing users to iteratively refine and execute evals in Harbor format. It maps agent components like prompts, models, and tools, and uses traces to identify patterns, enabling the creation of executable evals that can be tested against different models and versions. The skill supports continuous learning by transforming recurring user requests and errors into evals, facilitating agent improvement through iterative feedback loops.
Managed Deep Agents gives developers a managed way to build, run, and deploy Deep Agents with built-in runtime, streaming, sandboxes, evals, memory, and auth.

LangSmith Bring Your Own Cloud is now generally available on AWS, giving Enterprise teams managed observability, evaluation, and deployment inside their own VPC.

Learn what AI agents are, how they work in an LLM loop, and where workflows fit so you can build reliable, production-ready autonomous systems.
