Dynamic ranking of models on how well they orchestrate tools for real-world agentic tasks, based on signals like tool reliability, task completion, and steerability.
Oct 2, 2026
2,158,478 sessions
51 models
Pareto Frontier
Net Improvement x Cost/task from real Agent Mode tasks