
LangSmith benchmarks LLM architectures by evaluating performance and tracing, with the new langchain-benchmarks package enabling reproducibility and experimentation. The Q&A dataset tests RAG systems' ability to handle complex queries, highlighting differences in model and architecture performance. Results show that GPT-3.5 outperforms Mistral-7b in aggregate metrics, but the Mistral model's accuracy and hallucination avoidance are superior, emphasizing the importance of retrieval accuracy and system prompts.
Managed Deep Agents gives developers a managed way to build, run, and deploy Deep Agents with built-in runtime, streaming, sandboxes, evals, memory, and auth.

LangSmith Bring Your Own Cloud is now generally available on AWS, giving Enterprise teams managed observability, evaluation, and deployment inside their own VPC.

Learn what AI agents are, how they work in an LLM loop, and where workflows fit so you can build reliable, production-ready autonomous systems.
