
The experiment found that increasing context and tools leads to performance degradation, with o3-mini and claude-3.5-sonnet showing significant declines in performance compared to gpt-4o and llama-3.3-70B. Agents requiring longer trajectories degrade more rapidly, and irrelevant domains worsen performance for o3-mini, while claude-3.5-sonnet remains more stable. The study highlights that adding more domains and tools reduces performance, particularly for agents with longer trajectories.
Managed Deep Agents gives developers a managed way to build, run, and deploy Deep Agents with built-in runtime, streaming, sandboxes, evals, memory, and auth.

LangSmith Bring Your Own Cloud is now generally available on AWS, giving Enterprise teams managed observability, evaluation, and deployment inside their own VPC.

Learn what AI agents are, how they work in an LLM loop, and where workflows fit so you can build reliable, production-ready autonomous systems.
