Test Run Comparisons

Langchain··Submitted by Mads Kristian Nylund
AI ToolsAI GovernanceAI Ethics

Test Run Comparisons is a feature in LangChain that enables users to compare multiple test runs side-by-side, providing visual insights into inputs, outputs, and evaluation metrics. It allows for LLM-assisted evaluations and filters to focus on specific datapoints, such as correct or incorrect runs, to identify differences. The feature is designed to help developers debug and understand LLM performance by comparing test runs across iterations. LangSmith, the agent engineering platform, was initially limited in its test comparison capabilities but now offers a user-friendly interface for this purpose, with plans for private beta and broader rollout.

Read Article

More from Langchain

Related Articles