Datasets

Execution datasets. Not static files.

Static instruction datasets tell models what people think agents should do. Ours record what agents actually do in production — every trace, tool call, failure and recovery, collected with explicit opt-in.

Real operational behavior is the moat.

Reality

Reflects real operations

Execution data captures agents in production: real tool calls, real latencies, real cost curves and real failure modes. Synthetic data cannot replicate this.

Freshness

Updates with reality

The corpus grows as agents run. New models, new tools and new workflows show up in our datasets the moment they appear in production.

Moat

Hard to replicate

You cannot scrape this. The corpus compounds through a network of participating agents — the more agents that contribute, the more valuable the data becomes.

What's inside the corpus.

Every dataset is normalized, anonymized and versioned with full documentation.

Traces

Agent execution traces

Full step-by-step execution records: state, context, tool calls and outcomes.

Tool use

Tool call sequences

What tools get called, in what order, and which sequences succeed at scale.

Multi-agent

Multi-agent interactions

Coordinated runs across agent teams: handoffs, contention and delegation patterns.

Failures

Error & failure corpora

The rare failure data that in-house teams almost never collect: crashes, edge cases and boundary conditions.

Metrics

Execution metrics

Latency, cost and success-rate distributions tied to agent configurations and workloads.

Corrections

Human correction pairs

Mistakes paired with human interventions, ideal for reward modeling and fine-tuning.

Sample dataset catalogue.

Dataset Records Format Use case
AgentTrace v3.1 1.2M traces JSONL Training, RAG over tool use
ToolSeq v2 480K sequences JSONL Tool-call planning
FailureCorpus v1 96K failures JSONL + Parquet Recovery, robustness evals
MultiAgent v1 64K runs JSONL Coordination research
CorrectionPairs v2 310K pairs JSONL Reward modeling, fine-tuning

Licensing that fits your stage.

Research

Research licenses

Full access for academic and nonprofit research teams. Fast approvals, per-project terms, and citation credit.

Enterprise

Enterprise licenses

Bulk dataset access, private releases, priority access to new corpora and dedicated data support.

API

Streaming API access

Query datasets programmatically — streaming samples, filtered subsets and versioned snapshots through our API.

Want access to the execution corpus?

License proprietary execution datasets built from millions of anonymized real-world agent runs. Research and enterprise licenses available.