# Where these words came from Both lists are built by `build_words.py` from the two source files committed beside it. Nothing here is copied from the original game. ## Guesses — `guesses.json`, 11,846 words From [`wordnik/wordlist`](https://github.com/wordnik/wordlist), snapshot `wordlist-20210729.txt`, **MIT licensed**, filtered to five-letter lowercase ASCII. Wordnik publishes this list for word-game developers, which is exactly what it is being used for. ## Answers — `answers.json`, 4,603 words The **intersection** of the guess list with SCOWL's common-American tier (`wamerican` 2020.12.07, as shipped in `/usr/share/dict/american-english`), minus a short hand-curated block list. The intersection is the point. A word is a possible answer if it is a real headword (Wordnik) *and* common enough that a non-specialist has plausibly met it (SCOWL). That rule is stated, reproducible, and ours — as opposed to copying somebody else's editorial selection of which words are fair. SCOWL is distributed under a permissive licence requiring attribution, which `NOTICE` carries. ## What we deliberately did not use **The original game's 2,315-word answer list.** It is the product of a deliberate human curation pass, which is the strongest selection-originality argument of any list in this space and has never been litigated. We do not need it, so we do not ship it. We verified separately that the Wordnik list *contains* all 2,315 of those words — that is a statement about Wordnik's coverage, not a reason to redistribute the selection. Consequences worth stating plainly, since they affect every number on the site: - Our answer pool is **4,603**, roughly twice the original's. This game is **harder** than the original, and our measured solver averages are not comparable to published figures for it. - The famous results — SALET as the optimal opener, a 3.4212-guess mean, a proven worst case of five — are proven **for the original list**. They are cited on the site as exactly that, never as our own ceiling. Our own reference player opens **TARES** and averages **3.72**, and that number is computed by `solver.optimal_depth_report()` rather than quoted. ## Rebuilding ```bash uv run python envs/wordle_five/words/build_words.py ``` Deterministic and offline. The output must be byte-identical; if it is not, `CONFORMANCE.txt` changes too and the TypeScript port has to be re-verified.