AI Agent Evaluation & Reliability Datasets

Reproducible, schema-validated datasets and evaluation specifications to stress-test autonomous LLM agents: tool failure recovery, state consistency, and compliance reasoning.

100% synthetic 窶・no PII, no scraped content 28,922 rows validated, 100% pass SHA-256 provenance manifests Apache-2.0 free samples

1. What this infrastructure evaluates

Generic benchmarks measure clean execution paths. These datasets target the failure modes that actually break agents in production:

Failure modeWhat is testedDataset family
Silent state corruption Agent trusts stale state after a tool returns "success" with a no-op result 窶・belief revision and autonomous replanning required. MCP Agent Trajectory
Runtime exceptions 403 insufficient scope, 429 rate-limit backoff, resource lock conflicts 窶・recovery sequences, not just happy paths. MCP / Function Calling
Schema drift under multi-turn load Argument-level integrity across long runs: declared tools only, declared argument keys only, strict role alternation. Function Calling EN / JA
High-context Japanese tool routing Honorific variation, entity double-meaning, token-efficiency breakdowns specific to Japanese. Function Calling JA
Compliance reasoning Multi-jurisdictional step-by-step audit reasoning (financial compliance, data privacy) with grounded conclusions. Regulatory Compliance CoT

2. Measured validation evidence

Every number below is produced by running validate_jsonl.py (stdlib-only, included in the repository) against the shipped dataset files. No sampling 窶・100% of rows are validated. Last measured: 2026-09-27.

MetricResultMethodology
Schema validation pass rate 28,922 / 28,922 (100.00%) Deterministic structural validation of every row across all lots and editions
Invalid JSON rows 0 Strict JSONL parsing, UTF-8, one sample per row
Phantom tool calls / undeclared arguments 0 (fails the row) Every call checked against declared tool + argument schema
Near-duplicate rate 0 at threshold 3-gram Jaccard 竕・ 0.85 exclusion vs. all prior rows incl. shipped lots
Validator self-test 23 mutation classes Gold rows must pass; deliberately corrupted variants must be caught

Full per-lot breakdown: evaluation/EVALS.md ツキ Reproduce locally:

python validate_jsonl.py --track mcp_tool_use path/to/dataset.jsonl

Try before anything else 窶・free 50-row samples (Apache-2.0)

The free samples use the same schema as the full commercial datasets, so you can wire them into your evaluation harness (Axolotl, Unsloth, custom) first:

3. Free sample vs. commercial package

Free sample (Hugging Face)Standard 1KExtended 2K+
Rows50 per track1,000+ per track2,300+ per track, non-overlapping with 1K
LicenseApache-2.0Perpetual commercial license (non-redistributable)
Schema validation reportPublic (EVALS.md)Included per delivery (stats.json)
SHA-256 integrity manifest窶・/td>Included (SHA256SUMS.txt)
Overlap guarantee窶・/td>Hash + ID + 3-gram Jaccard audited against all prior lots
Enterprise invoicing / NDA窶・/td>Available on request

4. Packages & acquisition

Card settlement via Gumroad. Corporate invoicing, NDA, and wire transfer available by inquiry.

MCP Agent Trajectory

  • Multi-turn MCP tool-use trajectories
  • 403 / 429 / lock-conflict recovery
  • 1K standard ツキ 2K+ extended lots

Platform price 窶・see gateway

Acquire license

Function Calling (EN)

  • Argument-level correctness under changing state
  • Strict schema, error-recovery rows
  • 1K standard ツキ 2K+ extended lots

Platform price 窶・see gateway

Acquire license

Function Calling (JA)

  • Japanese high-context tool routing
  • Honorifics, entity ambiguity coverage
  • 1K standard ツキ 2K+ extended lots

Platform price 窶・see gateway

Acquire license

Regulatory Compliance CoT

  • Multi-jurisdictional audit reasoning
  • Grounded conclusions, cited provisions
  • 1K standard ツキ 2K+ extended lots

Platform price 窶・see gateway

Acquire license

Enterprise / Custom

  • Custom generation to your tool schemas
  • Multi-branch volume licensing
  • NDA, corporate invoicing, wire transfer

By quote

Submit procurement inquiry

5. Licensing & FAQ

License terms

Commercial packages grant perpetual organizational usage rights for model training, evaluation, and internal tooling. Standalone redistribution of raw data is prohibited. Free samples are Apache-2.0.

Is the data safe for commercial training?

Yes. All records are 100% synthetic 窶・no customer data, no scraped content, no PII, no copyrighted material. Each commercial delivery ships with a SHA-256 integrity manifest for audit trails.

How is quality verified?

A multi-stage pipeline: generation-time schema gating 竊・deterministic structural validation (declared tools/arguments only, strict role alternation) 竊・3-gram Jaccard 竕・ 0.85 dedup against all shipped lots 竊・independent LLM scoring 竊・final review. The deterministic stage is public and reproducible (evaluation/).

Can I evaluate before purchasing?

Yes 窶・run the free 50-row samples through your pipeline first. They share the exact schema of the full datasets.

Do new lots overlap with ones I already bought?

No. Every new lot is audited (hash, ID, 3-gram Jaccard) against previously shipped lots 窶・no budget is spent twice on the same row.