We wrote down what we believed, built the smallest thing that could prove us wrong, and let the evidence pick the next step.
Sometimes the evidence showed the game was wrong. Sometimes it showed the process was wrong.
DevSpark is a spec-driven way of building software: each change starts as a short written specification, goes through adversarial review before any code exists, and ends with the lasting record in code, tests and written knowledge, while the spec itself is temporary. ArrowSpark is the evidence. The repository is public, so every claim below links to the commit or pull request that shows it.
Rules before code
Belief
Write down what the game promises before building any of it.
Evidence
A constitution, a specification, a plan and two rounds of adversarial review landed before the first arrow. The first playable board arrived 3 h 14 min after the first commit, with twelve commits of governance and planning in front of it.
Next decision
Keep the gates in front of implementation, and keep the rules in plain classes that can be tested without a screen.
Hard is good. Unsolvable is not.
Belief
A difficult puzzle should never be able to become impossible.
Evidence
Removing an arrow only ever frees cells, so every legal move keeps a solvable board solvable. A short solver can prove each puzzle can be finished.
Next decision
No lives, no fail state, unlimited mistakes. A mistake costs score, never play.
Animation is not just polish
Belief
How an arrow leaves the board is presentation.
Evidence
Long arrows feeding out along their own bends, instead of sliding off rigidly, read as pulling a thread free, and became one of the most satisfying parts of the game.
Next decision
Let that image organize the design that followed, while the rules still remove an arrow the instant it is selected.
“Keyboard accessible” is not a plan
Belief
A requirement that Level Select works with a keyboard or gamepad is enough.
Evidence
The adversarial review read the template's code and found that opening the menu gave nothing focus, so a keyboard player would land nowhere. Nothing looked wrong on screen.
Next decision
Fix it before implementation, test it, and ask game-specific questions in review (who owns focus when this appears?).
Metrics would tell us what was hard
Belief
Deeper dependency chains, density and cascades would predict challenge.
Evidence
Six experiments hit every structural target exactly. In human play they rated about 2 out of 5 for challenge.
Next decision
Structural complexity is not perceptual complexity. Look at long, bent, interwoven arrows instead, and keep the metrics as descriptions only.
Contract before difficulty
Belief
The next step is harder puzzles.
Evidence
A board that looks overwhelming asks the player for trust. Harder puzzles needed a safety valve first.
Next decision
Show Me an Open Move: one legal arrow highlighted, never removed, costing five mistakes' worth of score. Then make it harder.
The screen should not constrain the puzzle
Belief
Puzzles have to fit the screen.
Evidence
Bigger boards made arrows smaller, which hurt reading, tracing and selecting all at once.
Next decision
A zoomable, pannable canvas before any large content. The puzzle defines the world; the screen is a window into it.
Experiments gave vocabulary, not a recipe
Belief
A set of experiments would add up to a formula for a good level.
Evidence
Six hand-built experiments taught useful concepts, but none was the level. The human evidence was one aggregate session, recorded per puzzle as unknown, not zero.
Next decision
Compose one real level by hand, judged by a person.
A rule you don't audit is a wish
Belief
We follow our own rule that durable files never point at temporary planning notes.
Evidence
An audit found review identifiers leaked into code comments, a test and the test README.
Next decision
Keep the explanation, remove the citation, and audit regularly.
Green checks were not enough
Belief
A spec is done when its checks are green.
Evidence
Spec 010's definition of done was a human answer. The designer played the Reference Knot to the end and said: “Yes, I would be happy to have a friend play this level.” One criterion failed against its original wording and was amended in the open; another was deferred with a written protocol.
Next decision
One tester is not outside validation. Build this public showcase to find out.
Verification can create scope
Belief
More verification always means more confidence.
Evidence
A developer readout was added only to gather closeout evidence, never served its purpose, and was removed in review. Review went through three revisions after the objective was already met.
Next decision
The convergence rule: a finding becomes work only if it blocks the objective. Everything else is classified and carried.
The designer's opinion is not enough
Belief
If the person who built the level likes it, it's good.
Evidence
One tester, who designed the level. No independent player had tried it before this showcase.
Next decision
You play it, and tell us what you think.
Only a finding that blocks the objective becomes more work. Everything else is classified (an accepted limitation, deferred work, or something learned) and carried forward with its reason.How the project learned this
Code, Tests, Knowledge.
Code says what the system does.
Tests prove the behaviours that matter.
Knowledge explains the current architecture, product rules and decisions.
Specs are temporary. Once a change ships, the lasting truth lives in these three places.
What worked. What didn’t.
What worked
- Keeping the rules in plain classes apart from presentation, so every rule is tested without a screen. Source
- Solver-backed content. Every shipped puzzle is proven solvable, with a mistake-free solution replayed in the tests. Source
- Adversarial review before code. It caught a Level Select plan that would have left keyboard players with nothing selected. Source
- Regression depth. Two required gates with over 4,000 passing checks run on every push. Source
- Treating specs as temporary and code, tests and knowledge as the durable record. Source
- Letting a human playtest overturn a green experiment instead of explaining it away. Source
- Recording checks that weren't run as not run. The final verify gate says exactly what its pass does not cover. Source
- Pull-request review after every gate was green, which still found a real bug (focus landing on a hidden entry). Source
What didn’t
- Structural metrics did not predict fun. Six experiments hit every target and rated about 2 out of 5 for challenge. Source
- The evidence base is one tester, who designed the level. No independent player had tried it before this showcase. Source
- A success criterion encoded the wrong assumption. It failed against its original wording and was amended in the open, with the original kept beside it. Source
- The fresh-player check was not run. It was deferred, with a written protocol, to this showcase. Source
- The project broke its own rule. Temporary review identifiers leaked into durable code comments, a test and the test README. Source
- Verification grew its own scope. A developer readout added to gather evidence never served its purpose and was removed in review. Source
- Durable knowledge went stale. Three statements contradicted the new code until pull-request review caught them. Source
- Early content was too simple. Most early puzzles were built from single-cell arrows a player can read at a glance. Source
The current showcase level.
- 46 × 32board
- 115arrows
- 1,348occupied cells (of 1,472)
- 92%density
- 207bends
- 43 cellslongest arrow
- 6legal opening moves
- 661dependency edges
These are structural facts about the board. They describe it; they do not judge it. The project learned that structural measurements do not say whether a puzzle is fun.Source
Now form your own opinion of the game.
Try the GameContinue with the Journey, the Evidence, orread the whole story.