Reproducible, schema-validated datasets and evaluation specifications to stress-test autonomous LLM agents: tool failure recovery, state consistency, and compliance reasoning.
100% synthetic 窶・no PII, no scraped content 28,922 rows validated, 100% pass SHA-256 provenance manifests Apache-2.0 free samples
Generic benchmarks measure clean execution paths. These datasets target the failure modes that actually break agents in production:
| Failure mode | What is tested | Dataset family |
|---|---|---|
| Silent state corruption | Agent trusts stale state after a tool returns "success" with a no-op result 窶・belief revision and autonomous replanning required. |
MCP Agent Trajectory |
| Runtime exceptions | 403 insufficient scope, 429 rate-limit backoff, resource lock conflicts 窶・recovery sequences, not just happy paths. | MCP / Function Calling |
| Schema drift under multi-turn load | Argument-level integrity across long runs: declared tools only, declared argument keys only, strict role alternation. | Function Calling EN / JA |
| High-context Japanese tool routing | Honorific variation, entity double-meaning, token-efficiency breakdowns specific to Japanese. | Function Calling JA |
| Compliance reasoning | Multi-jurisdictional step-by-step audit reasoning (financial compliance, data privacy) with grounded conclusions. | Regulatory Compliance CoT |
Every number below is produced by running validate_jsonl.py (stdlib-only, included in the repository) against the shipped dataset files. No sampling 窶・100% of rows are validated. Last measured: 2026-09-27.
| Metric | Result | Methodology |
|---|---|---|
| Schema validation pass rate | 28,922 / 28,922 (100.00%) | Deterministic structural validation of every row across all lots and editions |
| Invalid JSON rows | 0 | Strict JSONL parsing, UTF-8, one sample per row |
| Phantom tool calls / undeclared arguments | 0 (fails the row) | Every call checked against declared tool + argument schema |
| Near-duplicate rate | 0 at threshold | 3-gram Jaccard 竕・ 0.85 exclusion vs. all prior rows incl. shipped lots |
| Validator self-test | 23 mutation classes | Gold rows must pass; deliberately corrupted variants must be caught |
Full per-lot breakdown: evaluation/EVALS.md ツキ Reproduce locally:
python validate_jsonl.py --track mcp_tool_use path/to/dataset.jsonl
The free samples use the same schema as the full commercial datasets, so you can wire them into your evaluation harness (Axolotl, Unsloth, custom) first:
| Free sample (Hugging Face) | Standard 1K | Extended 2K+ | |
|---|---|---|---|
| Rows | 50 per track | 1,000+ per track | 2,300+ per track, non-overlapping with 1K |
| License | Apache-2.0 | Perpetual commercial license (non-redistributable) | |
| Schema validation report | Public (EVALS.md) | Included per delivery (stats.json) | |
| SHA-256 integrity manifest | 窶・/td> | Included (SHA256SUMS.txt) | |
| Overlap guarantee | 窶・/td> | Hash + ID + 3-gram Jaccard audited against all prior lots | |
| Enterprise invoicing / NDA | 窶・/td> | Available on request | |
Card settlement via Gumroad. Corporate invoicing, NDA, and wire transfer available by inquiry.
Platform price 窶・see gateway
Acquire licensePlatform price 窶・see gateway
Acquire licensePlatform price 窶・see gateway
Acquire licensePlatform price 窶・see gateway
Acquire licenseBy quote
Submit procurement inquiryCommercial packages grant perpetual organizational usage rights for model training, evaluation, and internal tooling. Standalone redistribution of raw data is prohibited. Free samples are Apache-2.0.
Yes. All records are 100% synthetic 窶・no customer data, no scraped content, no PII, no copyrighted material. Each commercial delivery ships with a SHA-256 integrity manifest for audit trails.
A multi-stage pipeline: generation-time schema gating 竊・deterministic structural validation (declared tools/arguments only, strict role alternation) 竊・3-gram Jaccard 竕・ 0.85 dedup against all shipped lots 竊・independent LLM scoring 竊・final review. The deterministic stage is public and reproducible (evaluation/).
Yes 窶・run the free 50-row samples through your pipeline first. They share the exact schema of the full datasets.
No. Every new lot is audited (hash, ID, 3-gram Jaccard) against previously shipped lots 窶・no budget is spent twice on the same row.