Research
Insights from millions of real executions.
We study how agents actually behave in production — across models, tools and workloads — and publish what we learn from the anonymized execution corpus.
Featured reports.
Each report is backed by the execution corpus and published with full methodology.
The State of Agent Execution
How success rates, latencies and costs vary across models, tools and workloads — from millions of anonymized runs.
Read reportThe Failure Divide
Why 43% of agent failures are recoverable, where they happen, and what separates teams that recover well from teams that don't.
Read reportThe Cost of Wrong Tools
How tool-selection errors compound into latency and cost, and the correction patterns that prevent them.
Read reportLatest findings from the corpus.
43%
of agent failures are recoverable without human intervention — when agents have access to recovery strategies learned from real executions.
2.4×
median latency increase on tasks where the wrong tool was selected at least once. Early tool choice dominates end-to-end performance.
6.1×
better recovery rates for agents trained on human-correction pairs compared with baseline fine-tuning on task data alone.
How we do research.
Reproducible, transparent and grounded in real executions.
Trace collection
Executions flow in from registered agents across production workloads and frameworks.
Anonymization
Personal data, payloads and identifiers are stripped before analysis.
Normalization
Heterogeneous traces become a consistent, queryable schema.
Analysis
Patterns, failures and metrics are studied across models, tools and workloads.
Publication
Findings, methodology and supporting data ship together, versioned and reproducible.
Responsible research
Our research follows three commitments:
- Only anonymized, consented data is analyzed.
- Findings are published with full methodology.
- No individual contributor is ever identified.
We publish negative results too — the corpus is most honest when nothing is hidden.
Stay ahead of the data.
Get our research reports and the underlying execution data. Follow StratScope for new findings from the agent execution network.
