The forcing functions
This is the fourth throughline from the Sharpee retrospective, and it is the evidence behind the defects table and the first post's claim that a consumer executing the real path was the instrument that kept up. As before, what follows is the analysis as the retrospective produced it, counter-evidence and verification intact, lightly edited for the blog and corrected where the adversarial pass found a claim overstated. The "I" in the body is the analyst that wrote it.
One correction applied up front: the original timeline called the engine stub "85-line" and said every Phase 4 test asserted against it. The verifier measured 84 lines and found that the message-framing unit tests touched no stub. The post says 84, and "every Phase 4 test that exercised the sandbox."
Thesis
Sharpee's platform was not designed and then populated with stories; it was excavated by them. From the first C# prototype in March 2023 through the Fernhill tutorial in July 2026, nearly every substantial platform capability landed within days of something built on top of the platform failing to compile, run, or read correctly, and the defects that mattered most were structurally invisible to the review practices that ran alongside. An action that validated but never mutated survived three complete refactors of all 43 standard-library actions in 24 hours and died the first time a story tried to put a cloak on a hook. A save-and-restore service that dropped the score ledger, all capabilities, all state values, and all relationships survived from inception through a 3,164-test audit and a static grader reporting 177 GREEN, 0 YELLOW, 0 RED, because the three tests covering it asserted on fields the broken code never touched. A story compiler shipped a full language surface before anyone discovered it could not declare a door. The claim holds: dog-fooding found defects design review did not. But the record also shows the inverse is false as a general law. A pure documentation exercise triggered the single largest cleanup in the repo's history with no story involved. And it shows two distinctive failure modes of the practice itself: forcing functions that never existed still shipped API into the platform, and the harness that made dog-fooding legible was itself silently skipping commands for months.
Timeline
2023-03-19. The first C# prototype already contains a Cloak of Darkness story file. The standard first test story for interactive fiction is present before there is a platform to test. Source: file modification times on the prototype, three projects, 18 files, 675 lines.
2025-07-02. The entire build system is thrashed in one day, CommonJS, file references, a shell script monkey-patching Node's module resolver, for the single purpose of making Cloak of Darkness load. The commit is 1,025 files and 300,159 lines added. Source: commit 331b0674, "Completed most refactoring and build is working."
2025-08-09. Three hours after seven phases were declared complete with "All builds: Passing with new structure," running Cloak of Darkness produces four defects at once: scenery unreachable as an indirect object, a custom HANG action never invoked, raw message IDs rendered instead of text, READ unimplemented. Source: the session summary of 2025-08-09.
2025-12-26. All 43 standard-library actions are refactored three separate times in one day: to a three-phase pattern, to a helper pattern, then to a four-phase pattern. Design review kills the middle abstraction twelve hours after 43 actions adopted it. Source: six commits.
2025-12-27, 05:12. Debugging Cloak of Darkness finds that the putting action never moved anything: it called the container's add-item method for validation and returned. The fix is a one-line move call. This survived the previous day's three full refactors of every action. Source: commit 29970c8f.
2025-12-27, 06:21. The looking action is found to have been checking a room property that does not exist, since inception. It is one of eight issues catalogued that morning, all found by running one story. Source: commit 9ee1ccca; the darkness-system issues document.
2025-12-27, 16:44. Project Dungeo launches with the word in its README: "Dungeo is the dog-fooding project for Sharpee." Its second deliverable is a gap analysis that enumerates 20 missing actions, 11 missing traits, 6 missing behaviors, and 6 missing systems before a single room exists. Source: commit b730bc87.
2025-12-27, 20:47. Three decision records (NPC system, daemons and fuses, combat) are written and implemented the same day Dungeo starts. Fourteen records are added in the five days from December 27 to 31. Source: commit 18115c57; git history of the records directory.
2026-01-05. The entity-centric action dispatch record, the capability-dispatch system still central to the platform, opens with a single sentence of motivation: when a player types "lower basket," the system routes to the wrong action. One unexpressible Zork puzzle produces the architecture. Source: ADR-090's context section.
2026-02-15. The dog-fooding harness is found to have been lying: the transcript tester silently skipped every command that carried no assertion. Adding real assertions to eleven walkthroughs takes 611 insertions and forces the removal of teleport and take shortcuts. Source: commit b68bb92e; the session summary of 2026-02-15.
2026-02-16. A story called Reflections is deleted in full, 58 files, two days after a full mirror-portal rewrite was designed for it. Two decision records continue to cite Reflections as their motivating context. Source: commit a9f20e90.
2026-03-23. Family Zoo tutorial development logs two issues, multi-word aliases don't resolve ("examine bush babies") and entity creation needs three or four calls repeated thirty times in one file, both stamped "Discovered during Family Zoo tutorial development." Source: the issues list at commit 69c77c90.
2026-03-26. The alias issue is root-caused and fixed platform-side: the command validator switches to maximal munch, with a 154-line resolution test. The same commit adds an open-inventory trait after an NPC inventory scope leak. Source: commit 2dfe4dd4.
2026-03-26. Counter-evidence. Writing a computer-science-foundations architecture document, no story, no runtime, surfaces eight issues in one sitting. Execution deletes the entity handler system (1,566 lines) and the legacy event-sequencing layer (1,208 lines) and drives 1,035 unsafe casts to zero over nine phases. Source: commit 369acf9f and its three follow-ups.
2026-04-23. A phase named "Deno Sandbox: Engine Subprocess and Turn Execution" had been declared complete with an 84-line echo stub standing in for the engine. It is discovered only when a live player on the public site sees "Waiting for the story to begin…". The remediation produces the No-Stub-Under-Test rule now in DevArch as rule 13a. Source: the session summary of 2026-04-23; commit 2364db85 makes the entry point production code the next day.
2026-04-28. The save-and-restore service is found partially implemented since inception, silently dropping the score ledger, capabilities, state values, relationships, and ID counters on every host. The original code carries the comment "Full deserialization would need to clear and recreate the world. For now, this restores entity traits and locations." Source: commit bf9b9564.
2026-06-22. A "naive reader" QA pass in a Docker container over the 31-chapter manual finds the published devkit crashing on every command; the book is un-followable from section 1.4. Root cause is the sibling publish tool rewriting imports to relative paths on publish; fixed by a config flip and bumping 28 packages. Source: GitHub issue #148; commit 36e8951d.
2026-06-23. Book QA files five platform defects in one day, each fixed with a regression test the same day: player unexaminable, handlers double-firing, missing Oxford "and," silent NPC movement, a literal placeholder in examine-self. Source: five issues, five commits.
2026-06-24. A decision record is written because a book claim was checked and found aspirational: the SCENERY entity type added no scenery trait, so scenery was takeable. The record quotes the book's own warning paragraph as the evidence and introduces a default-trait registry so the type keeps its documented promise by construction. Source: ADR-189's context section.
2026-07-15. A record generalizes dog-fooding into methodology: Chord becomes an "elegance oracle." Because Chord compiles down to the platform, any Chord form cleaner than hand-written TypeScript proves the TypeScript has a seam. The owner's framing: "making Chord cruder is not an option; that means something is inelegant in Sharpee." Source: ADR-222.
2026-07-17. The go-live gate record documents a live verification: "create the oak door" still throws a load error; the most basic furniture in interactive fiction fails to load. The Chord language had shipped regions, topics, pronouns, and phrasebooks first. Source: ADR-233's context section.
2026-07-18. Fernhill ships as the tutorial gate, and its commits carry the platform fixes it forced in the same diff: "Fernhill G3 phases 1-5 (254/254) + live derived state SHIPPED" and "dynamic channels SHIPPED + Fernhill Phase 8 browser media." The derived-state record's context reads "Found live during Fernhill Phase 4." Source: three commits; ADR-240.
2026-08-02. Running Fernhill under a pinned seed reveals that no host ever passed the seed option through to story creation; every Chord story's random draw was clock-seeded, and the IDE's play mode had never run at its own pinned seed. Four hosts required fixing. Source: commits 6cf5dbc2 and 6c0b27c6.
Analysis
The pattern predates the platform. A Cloak of Darkness story file sits in the first C# prototype from March 2023, in a project whose action bodies, take, examine, go, are all empty. The test story arrives before the thing it tests. That ordering is the whole argument in miniature, and it recurs at every scale afterward.
Cloak of Darkness: the highest defect-per-line yield in the corpus. Cloak is a few hundred lines of story. In July 2025 it is the sole reason the build system is torn apart and rebuilt; it does not merely test the build, it defines what "working" means. On 2025-08-09 it produces four defects in a platform declared complete three hours earlier. Its best find comes on 2025-12-27: the putting action had been calling the container's add-item method for validation and returning without ever moving the item. The previous 24 hours had migrated all 43 actions three separate times, plus produced a hand-written logic assessment of every one of them. Four passes of pure design review over the exact file. None saw that the action did nothing. One story did, immediately. The same morning, the looking action was found reading a property that does not exist, one of eight issues in the darkness-system document.
Dungeo is the only forcing function launched as one, and it is the only one whose gaps were enumerated in advance. Its README states the intent outright, and the gap analysis written the same hour lists 20 missing actions, 11 traits, 6 behaviors, 6 systems. That document is the thesis in artifact form: the gaps are legible only against a specific, complete, externally fixed target. Fourteen decision records land in the five days that follow, and three of them are written and implemented on day one. The mechanism stayed constant into January: the capability-dispatch system that still underpins the platform opens by describing "lower basket" routing to the wrong action. One unexpressible puzzle, one architecture.
Family Zoo yields less, and yields differently, which is informative. Its two findings are both stamped "Discovered during Family Zoo tutorial development." Neither is a correctness bug in the Dungeo sense; one is a parser resolution failure on multi-word aliases, one is pure developer experience. A tutorial exercises breadth of the authoring surface rather than depth of world logic, so it surfaces ergonomics and resolution edges. The alias fix changed the command validator and added a 154-line test; the verbosity issue was never fixed as specified. The verbose pattern shipped as-is, judged to "teach well."
The book is the highest-leverage forcing function the project found, because a reader cannot be stubbed. The Docker "naive reader" pass caught the published devkit crashing on every command, a defect not in Sharpee source at all but in the sibling publish pipeline, reachable only by installing the published thing and following the instructions. Five more platform defects were filed on 2026-06-23, all fixed the same day with regression tests. And the scenery record is the purest instance in the record: the book had to warn readers that the SCENERY type produces takeable scenery, and the record cites that paragraph as its reason to change the platform. Documenting a promise the code did not keep is what made the gap actionable.
By July 2026 the practice becomes doctrine. A record names Chord an "elegance oracle": because Chord compiles down to the platform, any Chord construct cleaner than hand-written TypeScript is proof the TypeScript has a seam. That converts dog-fooding from bug-finding into a quality metric with an argument behind it. The go-live record then turns "ready" into four assertable gates, and its context records the finding that justifies the whole exercise: after a language surface with regions, doors by name, topics, pronouns, comments, imports, counters, and phrasebooks, "a door" still threw a load error. Fernhill produced commits that contain the story and the platform fix in one diff, and what it found was a cache invalidated off an enumerated list of eleven event types, an enumeration that, as the record says, "can never be complete." No review finds that. Running a story that mutates state outside the list finds it on the first turn.
Where the thesis needs qualifying. Design review did produce the single largest structural cleanup in the repo: writing the foundations document surfaced eight issues in one sitting and led to deleting two whole layers. So the honest claim is narrower than "dog-fooding beats review." Review reliably finds structural problems: duplication, dead layers, weak typing. It reliably misses behavioral ones: code that runs and does the wrong thing. Every defect in the timeline above is behavioral. And the April 2026 stub episode shows the real variable is not "a story" but "the production path executed." That phase had a story, a test suite, and a stub, and stayed broken until a live player saw the failure. Dog-fooding works because it is hard to fake, not because stories are magic.
Carried forward to October 8. The mechanism repeated at December's tempo. A port of a 2009 Textfyre story, The Secret Letter, opened on 2026-08-21 and became the platform's QA instrument within days, filing 65 issues that were each worked around in content or escalated into a decision record; six records were written and shipped on its account, including a 285-file corpus cutover that removed it and create the player from the language. The port was tabled on 2026-09-08 and a sixteen-phase engine survey took its place, all sixteen phases executed in two days. In October a 60-room story was started room by room so the testing harness could be designed against an author actually writing, after the owner clicked through the app and called its three-column tab unreadable to anyone who did not grow up with it.
Counter-evidence
The retrospective argued against its own thesis before publishing it. These are its objections, kept whole.
- Design review found more, structurally, than any story did. Writing the foundations document, no runtime, no story, about 550 lines of prose, surfaced eight issues in one sitting, and their execution deleted the entity handler system, the legacy event-sequencing layer, and drove 1,035 unsafe casts to zero across nine phases. Nothing was being dog-fooded. The same shape appears on 2025-12-26 when a self-authored architectural assessment killed the helper abstraction twelve hours after all 43 actions had adopted it.
- Reflections is a forcing function that never ran, and it shipped API anyway. Four records cite it; one states "This became a concrete blocker during Reflections development." The August 2025 digest records the story as not started, and the February 2026 commit deletes all 58 files two days after a full rewrite was designed for it. The player-switching feature it motivated nevertheless shipped: the method exists in the engine, and at the retrospective's cutoff a repo-wide search returned exactly one occurrence, its own definition. Zero callers, thirty months in. Reversed since: on 2026-08-27 the language gained
change the player to, and the engine now calls the method at the turn boundary; three stories use it. The caller that finally arrived names the opening protagonist rather than switching mid-play as Reflections imagined, and the first thing it hit was a guard that would have refused every actor in every test engine. The API found a caller because a story needed to say who the player is. - The dog-fooding harness was itself silently broken for months. The transcript tester skipped every command that carried no assertion; "they appeared in the transcript but were silently skipped during execution." The fix had to add real assertions across eleven walkthroughs and rip out the teleport and take shortcuts. Removing those shortcuts a month later exposed two genuine Dungeo map bugs the shortcuts had been hiding. So an unknown fraction of Dungeo's "passing walkthroughs" during January and February 2026 proved nothing.
- The dog-fooded regression numbers do not reconcile. Per the July 2026 digest, Dungeo walkthrough-chain totals are reported as 885, 916, 870, 874, 888, 921, and 866 across the month, in a chain the project documents as deterministic at a pinned seed. Determinism only became real on 2026-08-02 when four module-scope random-number singletons were deleted.
- The forcing function generated its own load. One record is a 34-point puzzle, a hidden painting and a lethal temple ritual, invented from scratch inside a project whose premise was faithful reimplementation of mainframe Zork. Its stated context is "Dungeo needs additional puzzles to reach the 616-point target." The dog food was being manufactured to keep the exercise going.
- Dungeo's yield is now explicitly quarantined. Project memory records it as a nostalgia project, "never a consideration for Chord or Chord Writer, not as corpus stats, scale yardstick, or example." Fourteen records and roughly 40% of all record files mention it, but the current products treat its requirements as non-representative, meaning some share of what it forced into the platform was Zork-shaped rather than interactive-fiction-shaped.
- Absence of a forcing function was sometimes correctly detected without one. The context-menu record was implemented in full, seven phases, about 4,600 lines, working, and then deferred with the reason written into its status: "the feature needs a concrete game and UX vision to drive it." It sits unmerged. The project could recognize a missing forcing function by reasoning, but only after building the feature.
- Some of the worst defects were found by neither stories nor review. One record was authored, reviewed, and rejected on an incidence count showing the bug it fixed occurs zero times across every story and the platform defaults. The 1,191 type-only symbols in value-import position across 456 files came from a purpose-built repo-wide detector. Neither is a dog-fooding find.
- Having a story is not sufficient; executing the real path is the actual variable. The sandbox phase was declared complete while the engine entry point was an echo stub and every sandbox test asserted against a test fixture of it. Dungeo was the story being served. The failure surfaced only when a live player saw the empty state. This is why the remediation produced a methodology rule rather than a story.
- Dog-fooding did not prevent long-lived silent breakage even where stories exercised it daily. The save-and-restore service dropped the score ledger, capabilities, state values, relationships, and ID counters on every host from inception until 2026-04-28, behind a comment reading "For now, this restores entity traits and locations." Cloak, Dungeo, and the multi-user server all saved and restored throughout that period without surfacing it.
What was verified
Verified directly against the repository: the putting fix and its exact diff; the looking fix and the darkness document's eight issues; Dungeo's README sentence and the gap analysis's four tables (20, 11, 6, 6 counted from the file); the fourteen records added between December 27 and January 1; the "lower basket" context; the puzzle record's context sentence; the Reflections deletion's stat and file count; the player-switching search (one hit, its own definition); the Family Zoo issue stamps; the alias fix's message and stat; the foundations commit and the two large deletions; the save-and-restore fix and the pre-fix comment; the six book-QA issues with creation dates and the five fix commits; the scenery record's context; the go-live record's door sentence and its tutorial amendment; the oracle definition; the derived-state record's "Found live during Fernhill Phase 4"; the Fernhill commit titles; the seed fixes; the context-menu record's deferred status and surviving branch; the harness fix's message and stat; record counts by month; the Dungeo mention count.
Taken from the digests and not independently verified in this pass: the 2023 prototype file times and inventories; the claim that a module-type flag was added and removed from every package on 2025-07-02 hours apart; the 3,164-test inventory and the 177 GREEN result (both confirmed later by the adversarial pass); the live player seeing the empty state (digest citing the session summary, which this pass did not open); the July walkthrough-chain totals; the 1,035 cast count (the phase commits and two large deletions verified, not the total); Fernhill's pattern and transcript counts (from the record's own amendment, so repo-sourced but self-reported).
One digest citation is wrong and worth flagging: the mid-2025 digest attributes the working Cloak build to a commit dated 2025-07-04 with the message "Core is unit tested." The commit matching the described work is 331b0674 from 2025-07-02, which is cited instead.
Two gaps the analyst could not close. First, no attempt was made to quantify how many of the 129 Dungeo-mentioning records were driven by Dungeo versus merely using it as an example; "mentions" is a weak proxy and no percentage claim is built on it. Second, there is no clean counterfactual for the central claim: no record exists of a design review that examined the putting action or the save-and-restore service and passed them, only the strong circumstantial fact that four refactor-and-assessment passes over all 43 actions occurred in the 24 hours before the putting bug was found by running a story. That is suggestive, not dispositive, and the analysis says so.