Built with DevSpark

We wrote down what we believed, built the smallest thing that could prove us wrong, and let the evidence pick the next step.

Sometimes the evidence showed the game was wrong. Sometimes it showed the process was wrong.

DevSpark is a spec-driven way of building software: each change starts as a short written specification, goes through adversarial review before any code exists, and ends with the lasting record in code, tests and written knowledge, while the spec itself is temporary. ArrowSpark is the evidence. The repository is public, so every claim below links to the commit or pull request that shows it.

Belief Evidence Next decisionEvery beat below follows this order.
  1. Rules before code

    Belief

    Write down what the game promises before building any of it.

    Evidence

    A constitution, a specification, a plan and two rounds of adversarial review landed before the first arrow. The first playable board arrived 3 h 14 min after the first commit, with twelve commits of governance and planning in front of it.

    Next decision

    Keep the gates in front of implementation, and keep the rules in plain classes that can be tested without a screen.

  2. Hard is good. Unsolvable is not.

    Belief

    A difficult puzzle should never be able to become impossible.

    Evidence

    Removing an arrow only ever frees cells, so every legal move keeps a solvable board solvable. A short solver can prove each puzzle can be finished.

    Next decision

    No lives, no fail state, unlimited mistakes. A mistake costs score, never play.

  3. Animation is not just polish

    Belief

    How an arrow leaves the board is presentation.

    Evidence

    Long arrows feeding out along their own bends, instead of sliding off rigidly, read as pulling a thread free, and became one of the most satisfying parts of the game.

    Next decision

    Let that image organize the design that followed, while the rules still remove an arrow the instant it is selected.

  4. “Keyboard accessible” is not a plan

    Belief

    A requirement that Level Select works with a keyboard or gamepad is enough.

    Evidence

    The adversarial review read the template's code and found that opening the menu gave nothing focus, so a keyboard player would land nowhere. Nothing looked wrong on screen.

    Next decision

    Fix it before implementation, test it, and ask game-specific questions in review (who owns focus when this appears?).

  5. Metrics would tell us what was hard

    Belief

    Deeper dependency chains, density and cascades would predict challenge.

    Evidence

    Six experiments hit every structural target exactly. In human play they rated about 2 out of 5 for challenge.

    Next decision

    Structural complexity is not perceptual complexity. Look at long, bent, interwoven arrows instead, and keep the metrics as descriptions only.

  6. Contract before difficulty

    Belief

    The next step is harder puzzles.

    Evidence

    A board that looks overwhelming asks the player for trust. Harder puzzles needed a safety valve first.

    Next decision

    Show Me an Open Move: one legal arrow highlighted, never removed, costing five mistakes' worth of score. Then make it harder.

  7. The screen should not constrain the puzzle

    Belief

    Puzzles have to fit the screen.

    Evidence

    Bigger boards made arrows smaller, which hurt reading, tracing and selecting all at once.

    Next decision

    A zoomable, pannable canvas before any large content. The puzzle defines the world; the screen is a window into it.

  8. Experiments gave vocabulary, not a recipe

    Belief

    A set of experiments would add up to a formula for a good level.

    Evidence

    Six hand-built experiments taught useful concepts, but none was the level. The human evidence was one aggregate session, recorded per puzzle as unknown, not zero.

    Next decision

    Compose one real level by hand, judged by a person.

  9. A rule you don't audit is a wish

    Belief

    We follow our own rule that durable files never point at temporary planning notes.

    Evidence

    An audit found review identifiers leaked into code comments, a test and the test README.

    Next decision

    Keep the explanation, remove the citation, and audit regularly.

  10. Green checks were not enough

    Belief

    A spec is done when its checks are green.

    Evidence

    Spec 010's definition of done was a human answer. The designer played the Reference Knot to the end and said: “Yes, I would be happy to have a friend play this level.” One criterion failed against its original wording and was amended in the open; another was deferred with a written protocol.

    Next decision

    One tester is not outside validation. Build this public showcase to find out.

  11. Verification can create scope

    Belief

    More verification always means more confidence.

    Evidence

    A developer readout was added only to gather closeout evidence, never served its purpose, and was removed in review. Review went through three revisions after the objective was already met.

    Next decision

    The convergence rule: a finding becomes work only if it blocks the objective. Everything else is classified and carried.

  12. The designer's opinion is not enough

    Belief

    If the person who built the level likes it, it's good.

    Evidence

    One tester, who designed the level. No independent player had tried it before this showcase.

    Next decision

    You play it, and tell us what you think.

Convergence rule
A spec is complete when its objective is resolved, not when every possible observation has been eliminated.

Only a finding that blocks the objective becomes more work. Everything else is classified (an accepted limitation, deferred work, or something learned) and carried forward with its reason.How the project learned this

Durable truth

Code, Tests, Knowledge.

Code says what the system does.

Tests prove the behaviours that matter.

Knowledge explains the current architecture, product rules and decisions.

Specs are temporary. Once a change ships, the lasting truth lives in these three places.

Code at the top, tests and knowledge below, each joined to the othersCODETESTSKNOWLEDGEwhat it doeswhat is provenwhy it is this way
Worked and didn’t

What worked. What didn’t.

What worked

  • Keeping the rules in plain classes apart from presentation, so every rule is tested without a screen. Source
  • Solver-backed content. Every shipped puzzle is proven solvable, with a mistake-free solution replayed in the tests. Source
  • Adversarial review before code. It caught a Level Select plan that would have left keyboard players with nothing selected. Source
  • Regression depth. Two required gates with over 4,000 passing checks run on every push. Source
  • Treating specs as temporary and code, tests and knowledge as the durable record. Source
  • Letting a human playtest overturn a green experiment instead of explaining it away. Source
  • Recording checks that weren't run as not run. The final verify gate says exactly what its pass does not cover. Source
  • Pull-request review after every gate was green, which still found a real bug (focus landing on a hidden entry). Source

What didn’t

  • Structural metrics did not predict fun. Six experiments hit every target and rated about 2 out of 5 for challenge. Source
  • The evidence base is one tester, who designed the level. No independent player had tried it before this showcase. Source
  • A success criterion encoded the wrong assumption. It failed against its original wording and was amended in the open, with the original kept beside it. Source
  • The fresh-player check was not run. It was deferred, with a written protocol, to this showcase. Source
  • The project broke its own rule. Temporary review identifiers leaked into durable code comments, a test and the test README. Source
  • Verification grew its own scope. A developer readout added to gather evidence never served its purpose and was removed in review. Source
  • Durable knowledge went stale. Three statements contradicted the new code until pull-request review caught them. Source
  • Early content was too simple. Most early puzzles were built from single-cell arrows a player can read at a glance. Source
Reference Knot

The current showcase level.

  • 46 × 32board
  • 115arrows
  • 1,348occupied cells (of 1,472)
  • 92%density
  • 207bends
  • 43 cellslongest arrow
  • 6legal opening moves
  • 661dependency edges

These are structural facts about the board. They describe it; they do not judge it. The project learned that structural measurements do not say whether a puzzle is fun.Source

Now form your own opinion of the game.

Try the Game

Continue with the Journey, the Evidence, orread the whole story.

All optional

What did you think of the way it was built?

Did the story of how it was built make sense?
Did it change how you see the game?
Would you use a process like this?

Comments are anonymous and may be quoted publicly in the project write-up.