David Cornelson - GenAI Master

Tell the verifier to break it

The last four posts rest on a retrospective of the Sharpee project that a model wrote from a corpus of 1,533 session summaries and three git clones. A model-written history of a model-built project is exactly the kind of document nobody should trust on its face, and I didn't. This post is about what was done to it before any of it was published, and what that process turned up.

The instruction

Every load-bearing claim in the retrospective was pulled out into a batch and handed to a separate verifier, with this instruction: break it. Default to refuted when the evidence is not clean. Say unverifiable when no primary source settles it. Reproduce every number rather than accept it. Forty-one claims went through. Twenty came back confirmed. Nineteen came back partly wrong, with corrected wording. Two were refuted outright.

When the corpus was extended on October 9 to cover August 10 through October 8, the same instruction ran again over 127 claims from the carried-forward throughlines: 80 confirmed, 45 corrected, 2 refuted. The record now holds 168 verified claims.

The rule attached to the result was the important part: anything marked other than confirmed carries corrected wording, and that corrected wording, not the original claim, is what the retrospective is allowed to say. That rule is why the earlier posts open with corrections instead of burying them.

What "partly wrong" looked like

Most of the nineteen were not errors of fact so much as errors of precision that a hostile reader would have found in one command.

A claim said a fluent TypeScript authoring layer was "written the same morning" as the Chord language design. The verifier opened the file: "Status: Design exercise. No implementation." A hypothetical package name used throughout, never created. It was a paper sketch, and the claim had described it as built.

A claim said a story's TypeScript canon was "deleted nine days after" a decision record. The verifier ran the commit with rename detection: every file was a 100%-similarity rename into a new directory, and the commit message said so in its own words, "Nothing is deleted; every move is a tracked rename." All nine files are still tracked at HEAD. The claim was refuted.

A claim said 36 of 318 decision records carry a terminal status and six were deleted outright. The verifier parsed every status line: 24 unqualified terminal, 28 on a wider list, and the six "deleted" paths were five renames and one interim plan that was never a decision record. The verifier's note on that one is the kind of sentence you want from a verifier: "It reads as evidence of destroyed decision records, and it is the opposite."

A claim said detection latency for false completion claims fell from eighteen days and nine months in 2025 to 87 minutes in 2026. The verifier checked each number, found all four, and refused the sentence anyway: four anecdotes are not a latency distribution, and the nine-month case was introduced in 2025 but detected in 2026. "Do not say the latency fell."

A claim said a shipped application vendors a 25 MB runtime. The verifier found the 25 MB was the compressed tarball in the repo; the binary in the shipped bundle is 112 MB and the toolchain directory is 165 MB. A repo-size number wearing a bundle-size label.

The extension added one that cuts against the retrospective's own generosity. The original had excused five "no commits were made this session" summaries as a harness limitation, a commit agent's subprocess invisible to the writer. The verifier read the methodology's own rule and the commits: the writer runs before the commit and the summary is staged inside it, so the summary is stale the moment it lands, by design. Four of six commits checked carried a summary denying their own existence. The retrospective had been kind to the tooling for the wrong reason, and the tooling is mine.

What confirmed looked like

Confirmed did not mean "sounds right." It meant reproduced. The branch-tester count of 397 tests falling to 86 after a cutover was confirmed by checking out the parent commit into a detached worktree, running the suite, getting 397, then running it at HEAD and getting 86. The claim that 79% of live standard-library source files predate 2026 was confirmed by walking every file's first-add date with rename following, getting 78.6%, and then adding the caveat the claim lacked: that's the source directory only, and whole-package the figure drops to 70.7%.

The save-and-restore claim, that a broken service survived a 3,164-test audit and a static grader reporting 177 GREEN, was confirmed and then strengthened. The verifier reproduced the 177 by listing test files at the grader's commit, found the grader's single-directory yellow gate at line 92 and its "GREEN (default)" fallthrough at line 110, and then found something the retrospective had not: the April audit had read the specific test file, graded it "Excellent, one of the best test files in the suite," written the exact gap under its own Gaps heading, and recommended Keep. "That sentence is a better centerpiece than the grader's default-GREEN branch." It is, and it became one.

Things nobody asked about

Each verifier batch ended with a section the instruction didn't require, and it is the best material in the record.

The workflow guide that became DevArch's first document reached the DevArch repository at 10:25 on the morning of 2026-01-14 and was committed to the Sharpee repository it was written in at 18:09 that evening, eight hours later. The extraction was published before it was back-filled into its source. Nobody had claimed otherwise; the verifier just noticed the timestamps were in the wrong order for the story being told.

Claude attribution in commit trailers started on 2025-08-10 with an unversioned "Co-Authored-By: Claude," four and a half months before the first commit carrying a model name. A narrative that read the first named commit as "the day Claude arrived" would have been wrong by a season.

The project's instruction file did not grow monotonically. It peaked at 781 lines in April 2026 and was cut to 441 in one commit in May. Any throughline about instructions accumulating was contradicted by the data after April, and the verifier said so before anyone wrote one.

Two measurement conventions had to be stated or the numbers would not reproduce: the non-merge commit count is 2,042 on the main branch and 2,055 across all refs, and the prototype's file dates depend on whether build output is excluded. The verifier's instruction to the publisher: state the basis, or a reader who reruns the obvious command gets a different number.

Several commits were authored in a container at UTC while the rest of the repository is Chicago time, so day boundaries shift by one for anything committed after seven in the evening. A "nine days" claim sat exactly on that seam: eight days local, nine days UTC.

And one note about the process itself: the batch of claims handed to the verifier was in two places more wrong than the retrospective it was compressed from. The retrospective said the C# projects mapped "almost name for name"; the batch dropped "almost." The retrospective said the only surviving mutation emitter was a debug sink no shipping code constructs; the batch said no emission exists. "Whoever compressed throughlines.md into these claims introduced both errors; the source text is closer to correct." Summarizing a careful document is where the hedges fall off. The project had already named that failure, in a session two weeks earlier, as confidence laundering: a claim inherited across documents, each hop dropping the hedge and keeping the prose.

What this costs and what it buys

It is not free. Forty-one claims across five batches, and 127 more across eight in the extension, each verifier working in a fresh context against the repository and the corpus, reproducing numbers by running git and test suites rather than reading digests. The verifier notes alone run to a few thousand words.

What it buys is the right to publish. Nineteen of forty-one claims needed correction, and 45 of 127 in the extension. Some of those corrections reversed the meaning of the claim. If the retrospective had gone out as written, a reader with the repository and ten minutes could have found the renamed-not-deleted files, the paper-sketch-not-code, the 36-that-was-24, and every number in the document would have been suspect after that. The method does not make a model-written history true. It makes it checkable, and it puts the checking on the record next to the claims, so that a reader can see which numbers were reproduced and which were trusted from a digest.

That is also the argument of the first post, applied to prose instead of code. The unit of verification is the demonstration. For a claim about a repository, the demonstration is the command and its output, dated, in the document. A citation to a prior document's assertion is not evidence; it is the thing that needs evidence.

← All posts