Official AtCoder datasets · licensed for AI
Problem statements, full test data and millions of verified human solutions — straight from AtCoder, first-party and consent-based. Train on it, or build verifiers and contamination-aware evals with it.
The non-public judge data: full test sets, special-judge checkers and scoring config. Build reward signals and verifiers that actually verify.
Statements in Japanese and English, official editorials, and a reference solution — a self-contained benchmark per problem.
Millions of accepted submissions across 100+ languages — real, diverse, judge-verified. Not synthetic.
The failed attempts too, in submission order — how humans actually debug their way to a correct solution.
Every submission is flagged in-contest vs post-contest, so you know exactly what was written independently under contest conditions.
Optimization contests with scored solutions and standings — human trial-and-error on open-ended problems, found nowhere else.
Statements, test data, editorials, reference solutions. The verifier kit.
The problems plus 500 accepted solutions per problem & language. The recommended entry point.
Every verified-correct solution ever judged.
All of the above plus failed attempts — complete human trajectories.
Per-solution pricing with built-in volume discounts, scoped however you want — by problem, language, contest or series. Pricing is differential: anything your company already licensed is excluded automatically, so you never pay for the same data twice.
Full price list and interactive charts will be available at launch.
Join the waitlist