David Cornelson - GenAI Master

The defects, by what found them

This is the table the first post leans on. It needs two paragraphs of context.

Sharpee is an interactive fiction platform I have been building since 2023, first as C# prototypes and since mid-2025 as a TypeScript monorepo, almost entirely through Claude Code. In August 2026 I assembled every session summary the project ever produced, about 1,250 of them across three and a half years, into one corpus and ran a retrospective over it. The retrospective's load-bearing claims, 41 of them, were then handed to a verifier instructed to break them: 20 came back confirmed, 19 corrected, 2 refuted. The corrected wording is what appears here. On October 9 the corpus was extended with the 244 summaries written since August 10, each throughline was carried forward through October 8, and 127 more claims went through the same verification: 80 confirmed, 45 corrected, 2 refuted. Where the extension reversed an entry below, the entry says so.

Two of the retrospective's throughlines, "The Forcing Functions" and "The gap between done and true," walk through the defects that mattered. What follows is each of those defects classified by the channel that caught it, with what had already looked at the code and not seen it. Where the throughline didn't state the channel, I went back to the session summary for the day and settled it from there.

Channels

  • Consumer: a real path executed by something that could not be stubbed. A story, a live player, a reader following the manual, a seed-pinned run.
  • Mechanical: a test, build, type check, grep, or audit script produced the signal without a human reading the code first.
  • Reading: a human or model read source, a spec, or a document and saw the defect before anything executed.

Marks on each entry: [V] the retrospective verified it against git or the file; [D] it comes from a monthly digest and was not independently verified; [C] the channel was settled from the session summary rather than the throughline.

How to read "missed by"

It is not a verdict on those instruments. Nearly every review pass in this record is a model reading model-written code, in a timeline compressed past the pace of any human-paced instrument. July 2026 alone saw 391 commits, 86 decision records, and 16 releases from one developer. Human reading was sampled at near zero, not measured and found wanting. "Missed by" is better read as "what didn't need to catch it": in a world where the fix lands the same day the consumer finds it, the instruments that missed these were not failing at their job. Their job had changed.

Behavioral defects: code ran and did the wrong thing

2025-08-09. Four defects in a platform declared complete three hours earlier: scenery unreachable as an indirect object, a custom HANG action never invoked, raw message IDs rendered instead of text, READ unimplemented. [D]

  • Found by: Consumer. Running Cloak of Darkness.
  • Missed by: Phase 1 through 7 completion claims; "All builds: Passing."

2025-12-27. The putting action validated via ContainerBehavior.addItem and returned without ever moving the item. One-line fix. [V]

  • Found by: Consumer. Debugging Cloak of Darkness.
  • Missed by: Three full refactors of all 43 actions plus a written IF-logic assessment of each, all inside the prior 24 hours. The retrospective's caveat: no review is on file as having examined putting.ts specifically and passed it.

2025-12-27. The looking action checked roomTrait.requiresLight, a property that does not exist; one of eight darkness-system issues. [V]

  • Found by: Consumer. Running one story.
  • Missed by: The same four passes.

2026-01-05. "Lower basket" routed to the mirror-pole action instead of the basket elevator. Produced ADR-090, the capability-dispatch architecture still in use. [V]

  • Found by: Consumer. Playing a Dungeo puzzle.
  • Missed by: Nothing had tried to express it.

2026-02-15. The transcript tester silently appended a skip assertion to any command written without one and returned before executing it; disembark, board, launch, and down had never run. [V]

  • Found by: Mechanical, then reading. Twenty-six cascading walkthrough failures, then reading the parser.
  • Missed by: Months of "passing walkthroughs" in January and February that proved nothing.

2026-03. Removing the $teleport and $take shortcuts exposed two genuine Dungeo map bugs the shortcuts had hidden. [V]

  • Found by: Consumer. Walkthroughs on the real path.
  • Missed by: The walkthroughs as previously written.

2026-03-23. Multi-word aliases didn't resolve ("examine bush babies"). [V]

  • Found by: Consumer. Authoring the Family Zoo tutorial.
  • Missed by: The parser test suite.

2026-04-23. A phase named "Engine Subprocess and Turn Execution" recorded DONE while the engine entry point was an 84-line echo stub; every sandbox test drove a 92-line fixture of that stub. [V] [C]

  • Found by: Consumer. A live player on play.sharpee.net saw "Waiting for the story to begin…"
  • Missed by: The Phase 4 test suite, all green against the stub; the plan tracker. The carve-out was disclosed in the commit message and the session summary's own status read INCOMPLETE, so this was known and not seen, not hidden. DevArch rule 13a was written the same day.

2026-04-28. The save-restore service had dropped the score ledger, capabilities, state values, relationships, and ID counters on every save since 2025-07-22, behind the comment "For now, this restores entity traits and locations." [V] [C]

  • Found by: Reading. Read while scoping the in-game save wire bridge.
  • Missed by: Three green tests asserting the hook was called and the event emitted, never on the saved data. A 3,164-test audit that read that test file, named the gap in writing ("No test for save operation providing actual save data back"), graded it Excellent, and said Keep. A static grader reporting 177 GREEN, 0 YELLOW, 0 RED. Nine months of three consumers saving and restoring daily.

2026-06-22. The published devkit crashed on every command; the book was un-followable from section 1.4. Root cause in the sibling publish pipeline rewriting imports. [V]

  • Found by: Consumer. A "naive reader" in a Docker container following the manual against the published package.
  • Missed by: The entire test suite; the defect was not in Sharpee source.

2026-06-23. Five platform defects in one day: player unexaminable, handlers double-firing, missing Oxford "and", silent NPC movement, a literal {description} in examine-self. [V]

  • Found by: Consumer. Book QA.
  • Missed by: The test suite; each fixed with a regression test the same day.

2026-06-24. The SCENERY entity type added no scenery trait, so scenery was takeable; the book had to warn readers about it. [V]

  • Found by: Consumer and reading. A book claim checked against behavior.
  • Missed by: The type system's documented promise.

2026-07-17. "Create the oak door" threw a LoadError after Chord had shipped regions, topics, pronouns, and phrasebooks. [V]

  • Found by: Consumer. The go-live gate run live.
  • Missed by: The language surface as designed.

2026-07-18. A derived-state cache invalidated off an enumerated list of eleven event types, "an enumeration that can never be complete." [V]

  • Found by: Consumer. Fernhill Phase 4, the first turn that mutated state outside the list.
  • Missed by: Design review of the cache.

2026-07-24. Two plans recorded COMPLETE, "all 17 phases, each green before the next." Eighty-seven minutes later: the branch never built; two packages were absent from the build's package lists, so every real-path gate had been unrunnable. [V] [C]

  • Found by: Mechanical. A build, test, and real-path sweep.
  • Missed by: Seventeen phases of unit-test-only green.

2026-08-02. No host ever passed the story loader's seed option through; every Chord story's random draw was clock-seeded and the IDE had never run at its own pinned seed. Four hosts fixed. [V]

  • Found by: Consumer. Running Fernhill under a pinned seed.
  • Missed by: A walkthrough chain the project documented as deterministic, whose totals had not reconciled for a month.

2026-08-02. The skip marker skipped execution, not just assertion; the third distinct failure of the same instrument in six months, contradicting the governing decision record's own text. [V] [C]

  • Found by: Reading, then consumer. Read the runner while scoping a plan, then proved it with a turn-zero save.
  • Missed by: Two prior fixes of the same instrument; the decision record's review cycle.

2026-08-10. The public site served a stale build; the deploy reported success and changed nothing, because the service's working directory pointed at a different clone than the deploy script built in. [V]

  • Found by: Consumer. The live site was observed stale.
  • Missed by: The deploy script's own success report.

Structural defects: code was wrong in shape, not in behavior

2025-12-26. A helper abstraction adopted by all 43 actions was killed twelve hours later. [V]

  • Found by: Reading. A self-authored architectural assessment.
  • Missed by: The migration that had just applied it.

2026-03-26. Eight issues surfaced in one sitting; the cleanup deleted the entity handler system (1,566 lines), the legacy event-sequencing layer (1,208 lines), and drove 1,035 unsafe casts to zero. [V]

  • Found by: Reading. Writing an architecture document, no story, no runtime.
  • Missed by: Everything that had been running fine on top of them.

2026-07. 1,191 type-only symbols in value-import position across 456 files. [V]

  • Found by: Mechanical. A purpose-built repo-wide detector.
  • Missed by: The type checker; it compiled.

2026-07. A decision record authored, reviewed, and rejected because an incidence count showed the bug it fixed occurs zero times across every story and the platform defaults. [V]

  • Found by: Mechanical. An incidence count.
  • Missed by: The record's own review.

2026-02 to 2026-08. A player-switching method shipped for a story that was deleted before it ran; zero callers, thirty months on. [V]

  • Found by: Mechanical. A grep during the retrospective.
  • Missed by: Four decision records that cite the story.
  • Reversed since: on 2026-08-27 the language gained change the player to, and the engine now calls the method at the turn boundary; three stories use it. The call that arrived names the opening protagonist rather than switching mid-play as the deleted story imagined, and the first thing it hit was a guard that would have refused every actor in every test engine.

False completion claims: the instrument said done, and it was not

2025-08-10. "Phase 1: COMPLETE (53 actions), Pattern Consistency 100%." Eighteen days later, 13 of 43 directories actually carried the pattern. There were never 53 actions. [V] The retrospective reclassifies this as denominator drift rather than one refuted claim.

  • Found by: Reading. Counting the directories.
  • Missed by: The summary that wrote it.

2025-09-01. "About 98% complete, 110+ tests, all passing, none skipped." Two hours later: three failing golden tests, one skipped, a broken AGAIN command. [D]

  • Found by: Mechanical. Running the suite.
  • Missed by: The summary that wrote it.

2026-01-12. "650/650 points complete (100%)," with a version bump. Decoding the 1981 FORTRAN data file gave a maximum of 616 and three treasures that do not exist in the original. [V]

  • Found by: Reading. An external primary source.
  • Missed by: The scoring work and the celebration.

2026-02-10. "All 7 steps complete, ready for build and test." Three hours later the extension didn't compile and 349 tests failed. [D]

  • Found by: Mechanical. Build and test.
  • Missed by: The summary that wrote it.

2026-04-06. The test grader reported "177 GREEN, 0 YELLOW, 0 RED." Its RED check was two greps, its YELLOW check applied only to one directory, and everything else fell through to a line reading # --- GREEN (default) ---. [V]

  • Found by: Reading. The retrospective's author read the script four months later. The project never found this itself.
  • Missed by: The project, for four months; the save-restore case three weeks later that the grader had graded GREEN.

2026-08-06. A summary claimed "no commits made this session"; the next session's mechanical HEAD check found the commit. [V]

  • Found by: Mechanical. The pre-session audit.
  • Missed by: The summarizing session. The retrospective called this a harness limitation, a commit agent's subprocess invisible to the writer. The extension refuted that: DevArch's rule 18 runs the summary writer before the commit and stages the summary inside it, so a summary saying nothing was committed lands in the commit that proves otherwise, on every terminal write, by design. Four of six later commits checked show the same thing. A methodology defect, not an over-claim, and DevArch's to fix.

Tally

Behavioral defects, 18 entries: consumer 13, mechanical 2, reading 2, one mixed. Structural defects, 5 entries: reading 2, mechanical 3, consumer 0. False completion claims, 6 entries: mechanical 3, reading 3, consumer 0.

Two entries deserve their own sentence. The save-restore defect is the only behavioral defect in the record that a human caught by reading, and it was caught after tests, a graded audit, and nine months of daily demonstration had all missed it. And the grader defect was never caught by the project at all. The instrument that certified test quality was itself unread until an outsider read it four months later.

What the record supports, and what it does not

It supports this: at compressed cadence, behavioral defects were found by executing the real path with a consumer that could not be stubbed, and were fixed at the speed they were found. Machine-speed checks found contradictions between claims and state. Model self-review found structure and little else. The consumer was not the strongest instrument. It was the one that kept up.

It does not support any claim about human reading's yield, since almost no human reading was happening at that pace and the reading entries are nearly all model reads. It does not support a counterfactual about what a human-paced team would have caught. And it does not support a rate: the retrospective's own finding is that its per-month contradiction counts are an artifact of reading budget, and this is a hand classification of the defects the retrospective chose to verify, not a census.

One more caveat on causation. 2025 had no methodology at all; DevArch's first commit is January 2026. The shift from eighteen days and nine months to 87 minutes and one session spans four model generations and one developer's accumulated experience, not instrumentation alone.

← All posts