
LangSmith is a platform for evaluating and comparing fine-tuned open-source models, enabling developers to create and compare evaluation datasets efficiently. The study compared Llama2-7b and Llama2-13b models using different training data volumes, showing that the 13b model outperformed the 7b model in accuracy, highlighting the impact of training data size and quality. LangSmith was used to evaluate the models, comparing their outputs to known correct answers and using GPT-4 for assessments, demonstrating the potential of open-source models to match or exceed established models like GPT-3.5T.
Managed Deep Agents gives developers a managed way to build, run, and deploy Deep Agents with built-in runtime, streaming, sandboxes, evals, memory, and auth.

LangSmith Bring Your Own Cloud is now generally available on AWS, giving Enterprise teams managed observability, evaluation, and deployment inside their own VPC.

Learn what AI agents are, how they work in an LLM loop, and where workflows fit so you can build reliable, production-ready autonomous systems.
