Datasets
Execution datasets. Not static files.
Static instruction datasets tell models what people think agents should do. Ours record what agents actually do in production — every trace, tool call, failure and recovery, collected with explicit opt-in.
Real operational behavior is the moat.
Reflects real operations
Execution data captures agents in production: real tool calls, real latencies, real cost curves and real failure modes. Synthetic data cannot replicate this.
Updates with reality
The corpus grows as agents run. New models, new tools and new workflows show up in our datasets the moment they appear in production.
Hard to replicate
You cannot scrape this. The corpus compounds through a network of participating agents — the more agents that contribute, the more valuable the data becomes.
What's inside the corpus.
Every dataset is normalized, anonymized and versioned with full documentation.
Agent execution traces
Full step-by-step execution records: state, context, tool calls and outcomes.
Tool call sequences
What tools get called, in what order, and which sequences succeed at scale.
Multi-agent interactions
Coordinated runs across agent teams: handoffs, contention and delegation patterns.
Error & failure corpora
The rare failure data that in-house teams almost never collect: crashes, edge cases and boundary conditions.
Execution metrics
Latency, cost and success-rate distributions tied to agent configurations and workloads.
Human correction pairs
Mistakes paired with human interventions, ideal for reward modeling and fine-tuning.
Sample dataset catalogue.
| Dataset | Records | Format | Use case |
|---|---|---|---|
| AgentTrace v3.1 | 1.2M traces | JSONL | Training, RAG over tool use |
| ToolSeq v2 | 480K sequences | JSONL | Tool-call planning |
| FailureCorpus v1 | 96K failures | JSONL + Parquet | Recovery, robustness evals |
| MultiAgent v1 | 64K runs | JSONL | Coordination research |
| CorrectionPairs v2 | 310K pairs | JSONL | Reward modeling, fine-tuning |
Licensing that fits your stage.
Research licenses
Full access for academic and nonprofit research teams. Fast approvals, per-project terms, and citation credit.
Enterprise licenses
Bulk dataset access, private releases, priority access to new corpora and dedicated data support.
Streaming API access
Query datasets programmatically — streaming samples, filtered subsets and versioned snapshots through our API.
Want access to the execution corpus?
License proprietary execution datasets built from millions of anonymized real-world agent runs. Research and enterprise licenses available.
