# How the Azadira problems were made

This document records where the problems come from, how they were written, how they are
checked, and what public sources they lean on. It exists so that anyone can audit the bank
for originality and correctness.

## Summary

* Every problem statement, example, test and recipe was **written from scratch for this site**.
  Nothing was scraped or copied from commercial judge sites, and their names do not appear in
  the content (a build check greps for them).
* The *algorithms* are public knowledge: binary search, Dijkstra's algorithm, dynamic
  programming, union-find and so on are textbook material. Each concept links to public
  references (Wikipedia, cp-algorithms.com, the Python documentation) from the problem's
  "Learning goals" tab; see `curriculum/references.js`.
* The problems were drafted by AI writing agents (Claude), one agent per topic file, from a
  topic brief that lists the techniques to cover. Every draft was then validated automatically
  before it entered the bank (see "Validation" below), and defects found by the validator were
  fixed by hand or by the agent.

## Process

1. **Curriculum design.** The five phases and the concept list for each phase were fixed first
   (`curriculum/phases.js`). Phase tutorials were written next, each referencing the core problems
   that practise its concepts.
2. **Core set (82 problems).** One writing agent per phase wrote the hand-curated core problems
   in `curriculum/problems/phaseN.js`, ordered to follow the tutorial.
3. **Problem bank (1,020 problems in 34 files of 30).** Thirty-four topic files (`curriculum/src/phaseN/<topic>.py`)
   were assigned to writing agents. Each brief named 20-30 techniques to cover, a difficulty
   mix, and the rule that every problem must be a *different task* framed in its own scenario
   (receipts, sensors, schedules, grids, games...) rather than a parameter variant.
   The purpose of the size is that a competition cannot be won by memorising the bank.
   A second round added **hard problems to every topic** (`<topic>__hard.py`, 199 problems): one
   writing agent per topic, with a brief requiring a clear insight, large hidden tests on which the
   obvious brute force is at least 20× slower (phases 2-5), and a check of every reference solution
   against an independently written brute force on hundreds of random inputs. Each idea was checked
   against the whole bank before writing, and an independent review afterwards removed 16 problems
   whose core question already existed. The **Gauntlet** (`phase5/gauntlet__*.py`, 21 problems) was
   written to be hard for AI solvers and tested blind; see `docs/AI-RESISTANCE.md`.
4. **Compact authoring format.** Bank problems are written as Python dicts. The author writes
   the statement, a reference solution, one to three **hand-verified checks** (which must match
   the worked examples in the statement) and a list of further inputs. The build tool runs the
   reference solution to compute the expected outputs for those inputs, so large-input tests are
   never hand-computed.
5. **Validation** (`tools/build_bank.py`, `tools/validate.mjs`). For every problem:
   * schema and unique id across the whole site,
   * the reference solution must load and must agree with every hand-verified check
     (a disagreement means either the example or the solution is wrong and the file fails),
   * the starter code must **not** pass the tests,
   * every test call must finish quickly (the browser is about three times slower than CPython),
   * outputs must be plain literals that round-trip through `repr` (large outputs are replaced
     by a checksum so the bank stays small).
   Files that failed were corrected until the check passed. Examples of defects caught: a
   task-scheduler solution that miscounted trailing idle slots, an example tree drawn with a
   different shape than its level-order list, a wildcard-matching example whose stated answer was
   wrong, inputs too large to grade in time.
6. **Originality pass.** After drafting, the content was grepped for the names of commercial judge
   sites and for verbatim example strings that are well known from them; those examples were
   replaced with new inputs and the outputs recomputed from the reference solutions.
7. **Recipes instead of solutions.** The site never shows a reference solution. Each problem
   carries (or is being given) a *recipe*: approach, ordered steps, a walkthrough of the first
   example, complexity and pitfalls, without full code. Reference solutions are kept only in the
   authoring sources and are stripped from everything the browser downloads
   (`curriculum/bank/*.json` contain tests but no solutions).

## Public sources used for concepts

| Area | Reference |
|------|-----------|
| Python language & standard library | https://docs.python.org/3/ |
| Data structures (arrays, hash tables, stacks, queues, linked lists, heaps, tries) | Wikipedia articles linked per concept in `curriculum/references.js` |
| Sorting, binary search, two pointers, sliding window, prefix sums | Wikipedia; https://cp-algorithms.com |
| Trees, graphs, shortest paths, spanning trees, union-find, topological sort | Wikipedia; cp-algorithms |
| Dynamic programming, backtracking, greedy algorithms | Wikipedia |
| General algorithm background | Cormen, Leiserson, Rivest, Stein, *Introduction to Algorithms* (concepts only; no text reused) |

Classic named problems that appear in every algorithms course (for example "maximum subarray",
"edit distance", "N-queens", "0/1 knapsack") are included under their standard names because the
names are part of the shared vocabulary of the field; their statements, examples and tests here
are original.

## Counts

82 core problems and 1,240 bank problems (1,322 total), every one with a recipe. Run
`python3 tools/build_bank.py` and `node tools/validate.mjs` for the current numbers.

## Review passes after drafting

Besides the automatic validation, later passes fixed issues a reader would notice: a text editor's
`erase(0)` that deleted everything, an example explanation that contradicted itself, a recipe step
with a left-in self-correction, an incorrect pitfall about slicing with `k = 0`, a problem title used
twice, a modulus constant used but never defined, and generated graph inputs with out-of-range node
ids. Recipe writers ran each reference solution on the worked example before writing its walkthrough.
The bank is generated from `curriculum/src/`; the write-up above applies to every file there.

## Reporting a problem

If a statement is ambiguous, an example is wrong, or a problem looks too close to something
published elsewhere, open an issue naming the problem id (shown in the page URL) and it will
be rewritten or removed.
