What we can show, and what we can’t.
Automated evidence
Over 4,000 passing checks
4,109 PASS lines across the two required regression gates at the Spec 010 merge: 4,031 in the puzzle gate and 78 in the save and input gate, with every failure counter at zero.
Reproduce:
At 6e60e11: python tests/run_puzzle_regressions.py --godot <godot-4.4> | grep -c '^PASS:' and python tests/run_regressions.py --godot <godot-4.4> | grep -c '^PASS:'Solver validation
Every shipped puzzle is proven solvable and its mistake-free solution is replayed. Order independence is checked at every branching state, and at every sixth one for the densest board.
Real-scene completion
Every catalog puzzle is played to the end inside the real game scene in a headless run.
Viewport and presentation
Coordinate conversion, zoom, pan and Fit, departure geometry, layout and focus order.
Saves and input
Settings recovery, input remapping and the guarantee that starting a puzzle never resets saved progress.
Human evidence
This showcase exists to gather that evidence. The optional reactions on the Play page and at the end of the story are how it arrives.
Reference Puzzle design report · Fresh-player check deferred to this showcase (commit)
Things not claimed
- No claim that the analyzer's measurements say whether a puzzle is fun.
- No claim that the Reference Knot is hard for everyone, or good for everyone.
- No claim of support on phones, tablets or touch screens.
- No claim of tested gamepad play in the browser.
- No speed-up multiple. Elapsed time is not effort, and there is no stated baseline to compare against.
- No claim that DevSpark is finished or right in every case.
Known limitations
- Accepted limitationThe Web-readiness gate the project set itself (three levels worth keeping and an outside tester) was not met. This showcase is how the outside tester arrives. Source
- Accepted limitationMid-level, around ten moves can be obvious at once. Recorded at the boundary of excess, not redesigned. Source
- DeferredWhether a first-time player finds the ArrowSpark Levels group in Level Select unaided. Source
- Not performedA recorded desktop smoke run and physical keyboard and gamepad traversal for the Reference Knot release. Source
- Accepted limitationThe six earlier experiments were judged in one aggregate session; per-puzzle results are unknown, not zero. Source
The clock
Counted from the first commit to the merge of the tenth spec (6e60e11). Elapsed time is not effort: nights are included in the first figure, and commit timestamps miss work done before a session’s first commit, so the session figure is a floor, not a timesheet. There is no speed-up multiple, because there is no stated baseline to compare against.
| Figure | Value | Reproduce |
|---|---|---|
| First commit to the Spec 010 merge, elapsed (nights included) | 121 h 10 min | git log --reverse --format='%h %ad %s' --date=format:'%Y-%m-%d %H:%M' 6e60e11 | sed -n '1p;$p' |
| Commits | 91 | git rev-list --count 6e60e11 |
| Merged GitHub pull requests | 4 | git log --merges --oneline 6e60e11 | grep -c 'Merge pull request' |
| Work sessions (a new session after a gap over 90 minutes) | 17 | git log --reverse --format=%at 6e60e11, then split wherever consecutive commits are more than 5,400 seconds apart |
| Commit-bracketed activity across those sessions (a floor, not a timesheet) | about 22.9 h | sum of (last commit − first commit) per session from the same split |
| Commits that touch tests/ | 19 | git log --oneline 6e60e11 -- tests | wc -l |
| Lines of GDScript in scripts/ and scenes/ | 3,820 | git ls-tree -r --name-only 6e60e11 | grep -E '^(scripts|scenes)/.*\.gd$' | xargs -I{} git show 6e60e11:{} | wc -l |
| Lines of GDScript in tests/ | 5,340 | git ls-tree -r --name-only 6e60e11 | grep -E '^tests/.*\.gd$' | xargs -I{} git show 6e60e11:{} | wc -l |
The Reference Knot’s board
- 46 × 32board
- 115arrows
- 1,348occupied cells (of 1,472)
- 92%density
- 207bends
- 43 cellslongest arrow
- 6legal opening moves
- 661dependency edges
These are structural facts about the board. They describe it; they do not judge it. The project learned that structural measurements do not say whether a puzzle is fun.Source