How Similarweb Evaluates Agent Reports with LangSmith

Langchain··Submitted by Mads Kristian Nylund
AI ToolsAI InfrastructureAI Evaluation

Similarweb uses LangSmith to evaluate long-form agent research reports, combining deterministic checks for tool usage and structure with LLM-as-a-judge scoring for meaning and quality. The evaluation process links results to traces and A/B comparisons, allowing detailed inspection of updates. Calibrated rubrics ensure accurate assessment, with scores and comments guiding further analysis of each change. The workflow emphasizes inspectability and repeatability, using feedback and traces to validate updates based on evidence rather than just appearance.

Read Article

More from Langchain

Related Articles