The gap between done and true
This is the second throughline from the Sharpee retrospective, the one the first post draws on for its section on instruments. What follows is the analysis as the retrospective produced it, with its counter-evidence and its "what was verified" section intact, lightly edited for the blog and corrected where the adversarial verification pass found a claim overstated. The "I" in the analysis is the analyst that wrote the retrospective, working from the corpus and the git history; the "I" in this paragraph is me.
One correction worth naming up front, because the original thesis made it and the verifier refused it: the retrospective wanted to say detection latency fell from 2025 to 2026. The verifier's ruling was that four anecdotes are not a distribution, and that one of them, a nine-month defect, was introduced in 2025 but detected in 2026. So the claim below is stated as four cases, longest to shortest, and not as a trend.
Thesis
Across three and a half years, Sharpee's 113 recorded contradictions between session summaries and git contain no dishonesty and very little deliberate over-claiming. They are overwhelmingly two other things: bookkeeping drift (status lines, denominators, phase ledgers) and, the dangerous minority, structural blindness, where the instrument that certified "done" was constitutionally incapable of seeing the thing it certified. The project's response was unusually good: it named the failure class (No-Stub-Under-Test, 2026-04-23), shipped a general methodology rule the same day, and built real-path tests as a habit that spread measurably through commit vocabulary. But the honest answer to "did the rate fall" is that the corpus cannot support a rate claim in either direction. What demonstrably changed is not the rate of false claims but the mechanism of their detection. The four longest-to-shortest cases run nine months, eighteen days, 87 minutes, and one session. The project did not cure over-claiming. It industrialised catching it.
Timeline
2025-07-22. A spatial-index deserializer lands in the engine carrying its own confession in a comment: "Note: Full deserialization would need to clear and recreate the world. For now, this restores entity traits and locations." Every save on every host silently drops the score ledger, capabilities, world values, relationships, and ID counters from this day forward. Source: commit 63bd3b3e.
2025-08-10 to 2025-08-28. "Phase 1: COMPLETE (53 actions), Pattern Consistency: 100%" is disproved eighteen days later at 27.5%. Git at the last August commit shows 13 files containing the pattern across 43 action directories. There were never 53 actions. Source: commit 7e3b2c88; the two session summaries.
2025-09-01 22:46 to 2025-09-02 00:43. "About 98% complete, 110+ tests, all tests passing, no skipped tests" is contradicted about two hours later by three failing golden tests, a skipped golden test, and a broken AGAIN command. The optimistic summary was written first and never corrected. Source: the two session summaries.
2026-01-03 to 2026-01-12. "650/650 Points Complete (100%)," with a version bump to celebrate, is disproved by decoding the 1981 FORTRAN data file: the maximum score is 616, three treasures do not exist in the original, and two value fields are systematically swapped. Nineteen files, 850 insertions. Source: commit ca192b8b.
2026-01-14. DevArch's first commit. Everything before this date, all of 2025's contradictions, happened with no session audit, no Behavior Statements, no test grading, and no lifecycle methodology of any kind.
2026-02-15. The transcript tester's parser was silently appending a skip assertion to any command written without one, and the runner returned before executing skipped commands. Five lines deleted; a hard validation error added. Twenty-six cascading walkthrough failures traced to four commands never having executed. Source: commit b68bb92e.
2026-04-06. The test grader reports "177 GREEN, 0 YELLOW, 0 RED. Phases 1-5 cleaned all issues." The script scans every test file under packages and stories, but its RED checks are two greps (a tautological assertion and a stray console.log) and its YELLOW check applies only to one directory of golden tests. Every other file reaches the line # --- GREEN (default) ---. Source: commit 00bc714e, lines 92, 110, and 113 of the script.
2026-04-19 to 2026-04-23. A phase named "Deno Sandbox Integration: Engine Subprocess and Turn Execution" is declared COMPLETE "with one explicit carve-out." The engine entry point was an 84-line echo stub whose own header says "Phase 0 stub: echoes INIT to READY"; every Phase 4 test that exercised the sandbox drove a 92-line Node fixture standing in for it. Found when a live player saw "Waiting for the story to begin…". Source: commit ac523102; session summary of 2026-04-23.
2026-04-23. The failure class is named. A document titled stub-antipattern is written the same day, along with a 201-line acceptance test running 5 of 5 RED against the real runtime: "The system under test cannot be the thing you wrote to stand in for the system under test." Source: commit 5395eecd.
2026-04-23. DevArch v2.2.0 ships "Integration Reality" the same calendar day as the incident that produced it, about half an hour before the incident report was committed. That exact sentence is in DevArch's ruleset today as rule 13a. Source: devarch commit e8d177e.
2026-04-24 to 2026-04-25. The engine entry point is rewritten for real, gate 5 of 5 GREEN; the stub apparatus stripped entirely. "Real-path" enters the project's commit vocabulary and never leaves: zero occurrences before this date, then 6 in May, 17 in June, 36 in July, 26 in the first ten days of August. Source: commits 2364db85 and da8cf6df.
2026-04-28. Save and restore fixed nine months and six days after the defect was introduced. A 262-line serializer replaced; a 314-line round-trip test added. The three tests that had been green the whole time asserted that the save hook was called and that the completion event was emitted, never on what was saved. Source: commit bf9b9564; the pre-fix test file at lines 401 and 422.
2026-07-24, 02:00 to 03:27. Two plans recorded COMPLETE, "Phases executed: all 17," "each green before the next." Eighty-seven minutes later: "The one real defect: the branch never built." Two extension packages were absent from the build tool's package lists, so all that green was unit-test-only and every real-path gate was unrunnable. Source: commit 93da76aa; the two session summaries.
2026-08-01. "Confidence laundering" is named: a claim originated unchecked in one model's response, was inherited by a decision record, briefly re-hedged, then reverted to the confident form, each hop dropping the hedge and keeping the prose. The claim turned out to be true; the defect was stating as established what had only been repeated. Rule saved: evidence inline, command and result, dated, never a citation to a prior document's assertion. Source: the session summary of 2026-08-01.
2026-08-02. The skip marker is found to skip execution rather than assertion: the third distinct failure of the same instrument in six months (implicit skip in February, skip-whole-command in March, execute semantics in August), and it contradicted the governing decision record's own text. Four new tests. Source: commit 8bcc3207.
2026-08-10. The public site had been serving a build from before the latest decision; the deploy reported success and changed nothing, because the service's working directory pointed at a different clone than the deploy script built in. Fixed, and a guard added that compares the service's working directory to the tree about to be built and refuses: "the bug was the silence, not the path." Source: commit 9cca4695; the session summary of 2026-08-10.
Analysis
Classification first, because the framing matters. Of the 113 contradictions, none are dishonesty. Every single one was discovered, written down, and preserved by the project itself, usually by the next session, in a file committed to the repository. A corpus that contained a cover-up would not contain the cover-up's exposure. What it contains instead sorts into five classes, and they are wildly unequal in importance.
Roughly 25 are status-line and numbering lag: a decision record reading DRAFT while shipping across six packages, two different records with the same number and opposite status lines, a record dated four months in its own future. Roughly 40 are denominator drift: 43 versus 47 actions, 1,035 versus 200 plus 527 unsafe casts, 521 versus 423 IDE tests, walkthrough chain totals of 885, 916, 870, 888, 921, and 866 reported as though comparable. These are bookkeeping. They are noisy, they make every metric in the corpus untrustworthy, and they are not verification debt in the interesting sense.
About 22 are the model over-claiming: writing reported as working. "All 7 Steps Complete, Ready for Build + Test" against a session three hours later finding the extension didn't compile and 349 tests failing. "Status: COMPLETE" for Docker packaging while the same file states the build was never executed. Five separate August 2026 summaries asserting "no commits were made this session" while git holds the commits. That last one is not over-claiming, and it is not the harness artifact the retrospective first called it. The October extension traced it to DevArch's own rule 18: the summary writer runs before the commit and the summary is staged inside it, so a summary saying nothing was committed lands in the commit that proves otherwise, on every terminal write, by design.
The dangerous class is small: about 15 cases of structural blindness, where the test could not see what it asserted, plus 2 stub-under-test. These are not exaggeration. They are instruments that return GREEN by construction.
The centerpiece is the grader. On 2026-04-06 the project built a static test grader implementing its own RED/YELLOW/GREEN taxonomy and reported "177 GREEN, 0 YELLOW, 0 RED. Phases 1-5 cleaned all issues." Reading the script: it walks every test file under packages and stories. Its RED detection is two greps. Its YELLOW detection is gated behind a path match on one directory of golden tests. Everything else falls through to a line that literally reads # --- GREEN (default) ---. So "0 YELLOW across 177 files" means no file outside that one directory was ever eligible to be yellow. Three weeks later the save-and-restore investigation found the platform-operations test file asserting that the save hook had been called and that a completion event had been emitted, textbook YELLOW by the project's own rule ("asserts on event emission but not on the state change the event represents"). The grader had graded it GREEN. An instrument built to measure verification quality was itself unverified, and reported the absence of two typos as the presence of quality.
The verifier found something stronger than the grader's default branch, and it belongs here: the April audit did not merely miss the save-and-restore test. It read the file, graded it "Excellent, one of the best test files in the suite," and listed the exact gap under its own Gaps heading: "No test for save operation providing actual save data back (only tests that the hook is called)." Then it recommended Keep.
That is the whole throughline in one artifact, and the save-and-restore bug is its twin. The truth had been sitting in the source in plain English since 2025-07-22, "For now, this restores entity traits and locations," for nine months, while three green tests looked past it at metadata fields the broken code never touched. Nobody lied. Nothing lied. The instruments simply pointed elsewhere.
What the project did about it is genuinely good, and it happened fast. On 2026-04-23, the day the stub was found, three things landed: a 201-line acceptance test running 5 of 5 RED against the real runtime; a document diagnosing why the methodology permitted it, rule by rule ("rule 11 catches tautological assertions and mock-only verification. It does not catch tests that assert on real state mutation inside a stubbed subsystem"); and DevArch v2.2.0 shipping the Integration Reality rule. The sentence "The system under test cannot be the thing you wrote to stand in for the system under test" travelled from an incident report into a general methodology product in under 24 hours, and is verbatim in DevArch's ruleset today as rule 13a. The habit spread measurably: "real-path" appears in zero commit messages before 2026-04-24, then 6, 17, 36, and 26 across May through the first ten days of August.
Did the rate fall? The evidence does not support a clean claim, and the reason is methodological. The monthly digests yield 8, 11, 8, 10, 8, 8, 8, 9, 10, 7, 10, 8, 8 contradictions per month, across months holding 7 session summaries and 258. Density therefore ranges from 1.14 down to 0.031 and is almost perfectly anti-correlated with corpus size, which is the signature of a fixed per-month output budget, not a measured rate. Anyone reading "113 contradictions" as a quantity has read a reading protocol, not a project.
Two independent measures do move. Correction vocabulary in commit subjects ("never built," "was wrong," "contrary to," "reported success") rises from 2 in February 2026 to 10 in July. That is more catching, not more erring, though it is equally consistent with more surface area shipped per month. And the four detection cases, read longest to shortest: nine months for save and restore, eighteen days in August 2025, 87 minutes for the branch that never built, and one session for the commit the summary said was never made. The last is the clearest: a summary claimed no commits, and the following session's mechanical HEAD check caught it, which is instrumentation doing exactly what the 2025 sessions had no way to do.
But the last month is not clean, and that is the finding. August 2026, the most heavily instrumented month in the project's history, with pre-session audits, plan review, decision-record interviews, mutation verification, and real-path gates all firing, still produced eight contradictions, still shipped a deploy that reported success and changed nothing, and still found the skip marker skipping execution for the third time in six months after two prior discoveries of the same behavior. The defect in the instrument that measures done survived being named twice. That is the honest shape of this arc: the class was correctly diagnosed, the remedy was real and travelled beyond this repo, and the failures kept coming anyway. Smaller, faster-caught, and better-documented, but not fewer.
Carried forward to October 8. The structural-blindness class did not fade; it acquired a name in the project's own issue ledger, "a check structurally incapable of observing what broke," six instances across five sessions. On 2026-08-28 a session found that the root tsc --noEmit type-checked nothing, because the root config lists no files and plain tsc does not follow project references, and wrote that every prior "root tsc clean" evidence line had been empty evidence. A sibling defect had test:ci counting a package named NONEXISTENT as a pass. On 2026-09-10 the project stopped counting instances and audited the gates as a system, and part of the fix landed within a day. The instrument that measures done kept failing, and the project kept getting faster at noticing.
Counter-evidence
The retrospective argued against its own thesis before publishing it. These are its objections, kept whole.
- The 113 figure cannot bear the weight the throughline wants to put on it. Contradictions per month are 7 to 11 across months containing 7 to 258 session summaries; density is anti-correlated with corpus size. That is a fixed per-month reading budget, not a measured rate. No trend claim in either direction survives this.
- Not one of the 113 is dishonesty, and the framing "the gap between done and true" risks implying otherwise. Every contradiction was found and recorded by the project, usually within days, in files it chose to commit. The corpus is a confession, and confessions are evidence of honesty, not its absence.
- Most of the 113 are not verification debt at all. Roughly 65 of 113 are status-line lag and denominator drift. These make metrics untrustworthy, but no false "done" rests on them.
- Several entries counted as contradictions are the remedy visibly working. One summary corrects itself inside the same file; one keeps both a false claim and its retraction, which the reader flags as "the honest form"; one session's own finding was disproved by a git search and retracted in three places including a public comment. Self-correction being logged as contradiction inflates the count against the project.
- The five August 2026 "no commits were made this session" claims are not over-claiming, but the retrospective's first explanation, a commit agent's subprocess invisible to the writer, was wrong. The extension found the mechanism in DevArch's rule 18 ordering: the summary is written, then committed, so it is stale the moment it lands. Four of six later commits checked carry a summary that denies their own existence. Counting them as verification debt still conflates a methodology defect with a false completion claim, but the defect is now DevArch's to fix.
- One decision record's current status line is a model of the discipline the throughline says was missing: "SUPERSEDED IN PLACE. Do not implement. Nothing here was built." By August 2026 the project writes negative results into status lines explicitly, which is the opposite of the "Proposed means shipped" pattern the retrospective warns about.
- A rising correction-commit count is ambiguous. It is consistent with better detection, but equally consistent with a project shipping more surface area per month. Commit volume roughly tripled over the same window, so the rate per commit may be flat.
- 2025 had no methodology at all; DevArch's first commit is 2026-01-14. Comparing 2025's eighteen-day detection against 2026's 87 minutes credits DevArch for a difference that also spans four model generations, a monorepo that grew tenfold, and a solo developer's own accumulated experience. The causal attribution to instrumentation is plausible, not demonstrated.
What was verified
Verified directly in git: the 13-of-43 count at the last August 2025 commit; the engine entry point at 84 lines and the test fixture at 92 lines at the phase-complete commit, both stripped two days later; the full text of the stub-antipattern document and its verbatim sentence in DevArch's ruleset; the DevArch commit dated 2026-04-23 titled "v2.2.0: Integration Reality"; the pre-fix save-and-restore comment and its origin in the engine on 2025-07-22; the pre-fix assertions in the platform-operations test at lines 401 and 422; the 262-line serializer rewrite plus the 314-line round-trip test; the complete source of the grader including the find at line 113, the single-directory YELLOW gate at line 92, and "GREEN (default)" at line 110; the five-line deletion of the implicit-skip default in the parser; the commit messages for the skip fix, the never-built branch, the score decode, and the documentation commit; the timestamp of the commit a summary said was never made; the "real-path" commit-message counts; DevArch's first commit date. The analyst also read the primary text of the deploy session and the confidence-laundering session rather than relying on a digest's paraphrase, and located the "177 GREEN" line in the correct session file after finding the digest had cited the wrong one.
Not verified: the underlying read of 1,247 session summaries (the digests are trusted for it); the 2025 timestamp pairs used for the eighteen-day figure, which come from digest-quoted summary headers; the FORTRAN decode itself; the 26-cascading-failure count. The classification counts (about 25 status lag, 40 denominator drift, 22 over-claim, 15 structural, 2 stub) are the analyst's own reading of the 113 digest entries, not an independent re-derivation from the corpus, and the boundaries between classes are judgment calls. No test suite or build was run.
The adversarial verification pass afterward confirmed the save-and-restore claim, the 3,164-test and 177-GREEN claim, the implicit-skip claim, the No-Stub-Under-Test promotion, and the reading-budget finding; corrected the stub claim (the message-framing unit tests touched no stub, and the DONE status was in the tracker while the summary's own status read INCOMPLETE and the commit disclosed the carve-out); corrected the 53-actions claim (two different measurements, not one refuted claim); and refused the latency-fell claim as stated above.