I’m the creator of Failure Map, an archive of compact Python debugging tasks. The open release has 20,168 tasks across 254 categories, with standard-library implementations, explicit contracts, failed repair attempts, and executable boundary checks. Three small cases to try: • Duplicate delivery: deduplicating equal amounts loses legitimate events. https://failuremap.org/cases/FA-001 • Cache expiry: subtracting a whole tick rejects an entry that is still valid. https://failuremap.org/cases/FA-006 • Pagination: changing > to >= repeats the cursor record. https://failuremap.org/cases/FA-011 Prompt template: “Repair the solve function to satisfy the stated contract. Return Python source only. Preserve the signature. Contract: {prompt}. Broken implementation: {broken_source}.” Measured program baselines, passed checks out of 3 (broken / attempted repair): FA-001 2/3 / 1/3; FA-006 2/3 / 2/3; FA-011 2/3 / 1/3. These are executions of the included programs, not model scores. I have no measured local-model results to claim yet. To compare runs, report the exact model and revision, quantization, prompt, sampling settings, seed, attempts per task, and pass counts. Run candidate code in isolation and keep grading fixtures outside its control. Recorded-check success is not hidden-test performance. Download: https://failuremap.org/api/exports/tasks.jsonl.gz Methodology: https://failuremap.org/methodology   submitted by   /u/failuremap-f [link]   [comments]