David Cornelson - GenAI Master

Forty-one weeks

The previous posts in this series use Sharpee's record as evidence for an argument about verification. This one is the picture the argument sits in front of. I want to put the velocity down in one place, with its receipts, because the earlier posts quote it in fragments and a reader who has not seen the whole shape will underweight what the fragments are describing.

Where it started

I designed Sharpee in C# in March and April of 2023. The design was finished: a world model, a data store, a grammar library, a parser library, a standard library, a text service. The decomposition in that prototype is recognizably the shape of the platform that exists now. Every action body was empty. I could not build it. The retrospective's testimony section has my own words for that year: the salad days of hallucinations and imaginary code.

The project was put down for eleven months. Picked up in 2024 and built in C# through 124 Claude conversations. Evaluated for a move to TypeScript in December 2024. First committed in March 2025, then silent in git for twelve weeks while I fought a build. Put down again for sixteen weeks in the autumn of 2025, with no commits, no summaries, and no conversations in any source.

It came back on Christmas Day 2025. The first commit of the return carries the first model-name attribution in the repository, and twenty-six minutes later the project's instruction file gained its first process rule, which was about running out of context. From that commit to October 8, 2026, where the record this series rests on closes, is forty-one weeks to the day. The 5.0.0 publish on August 10 sits thirty-three weeks in.

Where the record closes

On October 8, 2026, the project had:

  • 33 packages on npm at version 5.4.1, up from a first beta publish of 11 packages on January 1. The 5.0.0 publish was August 10; nine more followed in the eight weeks after it.
  • A story language, Chord, at 3.6.0 after a run of six minor versions and one rolled-back major, with a lexer, an indentation-aware parser, a two-pass analyzer, six load-time gates, a versioned intermediate representation, and a story loader. The platform's own English grammar is written in it: 488 lines of Chord generate the parser's grammar file, with a build gate that fails on drift.
  • A native macOS authoring app, Chord Writer, at 1.4.0 with auto-update: 187 Swift files, 32,613 lines, 78 test files, 605 tests, a sealed toolchain with a vendored Node runtime so an author never sees npm, signed with a curated five-key entitlement set. On the night of the 5.0.0 publish it had 158 files and 480 tests and was sitting in Apple's notarization queue; it shipped three days later.
  • A website with about 175 pages of documentation, up from 151 at the publish, and a browser playground with a Chord editor.
  • A 31-chapter, eight-volume author manual, written in five days in June, whose QA pass became the best test harness the platform ever had.
  • A roughly 191-room implementation of mainframe Zork as the standing regression baseline, with a walkthrough chain in the high 800s of assertions.
  • 369 architecture decision records, 318 of them at the publish.
  • 2,606 commits on the main branch, 2,495 non-merge across all branches.
  • About 1,500 session summaries, the corpus this series rests on. The retrospective was built from the 1,250 that existed at the publish; the 244 written since were added on October 9 and verified against it.

Three years and seven months from "We're going to design a new parser-based Interactive Fiction platform using C#" to a platform at 5.4.1 and a shipped, auto-updating macOS app. Forty-one weeks for nearly all of the building.

The publish was not an ending. The notarization queue cleared on August 13, and the eight weeks from the publish to the close of the record carried nine more platform releases, six Chord minors and one rolled-back major, and the app's climb to 1.4.0. The repository had commits on 57 of those 60 days, with no two consecutive silent days, while three new model lines arrived and were adopted without a pause.

The shape of it, by month

Commits per month, all branches, from the retrospective's git index, rebuilt October 9 to run to the end of the record. The 5.0.0 publish falls on August 10, with 198 of August's 408 commits before it:

2025-03      2  █
2025-06      1  █
2025-07     13  ██
2025-08     88  ████████████
2025-09      3  █
2025-12    155  █████████████████████
2026-01    452  ████████████████████████████████████████████████████████████
2026-02    129  █████████████████
2026-03     92  ████████████
2026-04    251  █████████████████████████████████
2026-05    105  ██████████████
2026-06    290  ██████████████████████████████████████
2026-07    405  ██████████████████████████████████████████████████████
2026-08    408  ██████████████████████████████████████████████████████
2026-09    230  ███████████████████████████████
2026-10     14  ██   (eight days, through October 8)

October and November 2025 are absent because nothing happened in them. The two low months in 2026, February and March, are not slow months. February's commit count is small because one housekeeping commit removed 275,597 lines; March is the month a documentation exercise drove 1,035 unsafe casts to zero across nine phases and deleted two whole layers of the platform.

The compressions

These are the intervals that make the monthly picture legible. Every one is verified against git in the retrospective.

A language in a day. On July 10 a session opened framed as "a small rathole," with "no implementation planned." The decision record for Chord was accepted at 11:46. The compiler landed at 17:30, five hours and forty-four minutes later: 57 files, 7,579 lines, 44 tests, and a frozen golden transcript suite. By 22:01 the same night, Cloak of Darkness was authored as a 101-line .story file and ran through an 81-of-81 gate. The npm release carrying the language shipped four days later.

A grammar rewritten in 21 hours. On July 24 one forum post from one person observed that the library wasn't available in readable Chord form. A first reading of that complaint shipped on July 25 and was superseded the same day by a record whose session opened "Started as 'ADR-265 is completely wrong.'" That record's six children landed between the evening of the 25th and the afternoon of the 26th. They removed a priority scheme from 382 call sites, inverted the platform's grammar so that Chord is the source and TypeScript is generated, and added extend and remove operations to the language. Sharpee 4.0.0 shipped the same day.

A month that would have been a year. July 2026: 391 non-merge commits, 86 decision-record documents, sixteen npm releases from 1.5.0 to 4.3.0. The IDE, dormant for 37 days while all of that landed, was found broken rather than stale on July 23 and rebuilt as a Chord authoring environment in four days.

A subsystem in an evening, gone in two days. A play-tree testing subsystem for the IDE: record accepted at 17:37 on August 3, nine-phase plan at 18:29, all nine phases landed by 01:12 the next morning. Seven hours and thirty-five minutes. Superseded thirteen hours later after I asked whether we wanted it or a transcript editor, and deleted on August 6: exactly 2,768 Swift lines and fifteen test files.

A server in nine days. Scaffolded April 19, Docker-deployed with acceptance criteria closed April 24, browser end-to-end tests added at 21:19 on April 28, deleted at 21:21. 268 files, 51,555 lines. The second attempt at the same problem went from an accepted record to six complete phases in about 27 hours in May.

Forty-three actions, three architectures, thirteen hours. December 26, 2025: every standard-library action migrated to one pattern by 04:04, to a second by 05:12, a written assessment killing the second at 05:15, and a third pattern replacing both by 17:12.

A manual in five days. June 20 to 25: 31 chapters, eight volumes, 7,897 lines.

A migration by fan-out. June 26: one insight about where text collapses to a string too early folded six planned decision records into one, and 418 templates and 236 call sites were migrated by a 99-agent parallel workflow in a session.

What the velocity cost

The earlier posts carry this in detail, so here it is in one paragraph. About a third of everything ever written is gone from the tree. Nineteen of the retrospective's forty-one load-bearing claims needed correction before publication, and when the record was extended to October, 45 of 127 more did. A save-and-restore defect lived nine months behind green tests. The dog-fooding harness silently skipped commands for two months and the same instrument failed two more times after that. Seventeen phases were recorded complete on a branch that had never built. Status lines on decision records stopped being load-bearing. The residue of fast deletion, the parts bins and the unmerged branches and the repudiated helper that still ships, is the specific waste the retrospective found. And on the night of the 5.0.0 publish, the app that was the point of the previous month had not shipped, because Apple's notary service crashed on three consecutive submissions. It shipped three days later, once a byte-identical archive that was accepted in 72 seconds on one submission and stuck for ten hours on another had killed every theory that blamed the archive.

What the velocity is not

The retrospective argued against its own velocity story, and the arguments hold.

The return date fits a holiday as well as a model release. Opus 4.5 shipped in late November; the return was Christmas Day. Two earlier model releases landed inside the sixteen-week wall and produced nothing.

The January surge is over-determined. On December 30 an esbuild bundle took the platform's load time from 81 seconds to 142 milliseconds, a 578-fold faster inner loop. Dungeo launched on December 27 as an explicit dog-fooding target with a 191-room map to fill. The model is one of three sufficient causes, and the data cannot separate them.

July's 391 commits against June's 290 cannot be attributed to Chord without noting that a new model line's first commit in the repository is July 2, eight days before the language was designed. The four-days-to-npm figure may measure the assistant as much as the language.

Attribution is a convention, not instrumentation. 194 commits carry a bare "Claude" trailer and 114 carry none; the convention only started in August 2025, and subagent work is invisible to it. Every per-model count measures what a commit script wrote.

And the larger context windows did not produce larger commits. What they produced was longer coherent sittings: more commits per active day, not more files per commit. That is the only capability claim the commit graph supports on its own.

The mirror

The numbers above are the record. This section holds them up against what we know about how fast humans write software, at what complexity, and at what quality. The benchmarks are from general knowledge and the arithmetic is mine; where it is an estimate, it says so.

Before times With a GenAI harness Calendar time to a platform and a language Design to release, COCOMO for the platform alone TADS 3, one author about 8 years COCOMO estimate, team of ten about 5 years Inform 7, one author, after Inform 6 about 3 years Sharpee, one person 41 weeks Non-test lines of code per developer-day Delivered lines, sustained for the length of the project Brooks, a system programmer about 10 McConnell, lifecycle range 10 to 50 A strong developer, best stretch 100 to 200 Sharpee, one person about 1,000 Benchmarks from the general literature; Sharpee from the October 8 tree, about 1,600 a day with tests. The confounders are in the text below.

Raw output. The literature on human coding rates is old and stable. Brooks put a system programmer at about ten delivered lines a day. McConnell's range across the whole lifecycle is ten to fifty. A strong developer on greenfield work with no meetings might sustain one or two hundred for a stretch, and nobody sustains that for a year. Sharpee's tree at the October 8 commit holds 586,346 lines of TypeScript and Swift across 3,612 files, counted directly, excluding build output and type declarations. 233,231 of those are tests. 353,115 are not: 204,495 in the platform packages, 68,050 in stories and tutorials, 32,613 in the Swift IDE, and the rest in tools, documentation snippets, and the website. Produced in about twelve active months once the dead months are removed, that is about 1,600 lines a day including tests, 1,000 a day without, every day, for a year. Against the ten-to-fifty benchmark that is twenty to a hundred and sixty times. Against the best-day number for a strong developer it is still five to ten times, sustained for a year.

Scale, the way estimators measure it. Put the 353 KLOC of non-test source into basic COCOMO as an organic project and it returns about 1,140 person-months, call it 95 person-years. Take only the 204 KLOC of platform packages, leaving out the stories, the IDE, and the tooling, and it is still about 53 person-years. The model is from 1981 and it is crude, but it is what every planning spreadsheet of the before times descended from, and it says the platform alone is a product a company would have staffed with a team of ten for five years.

The domain comparators, which are the fairer mirror. Interactive fiction has two solo-author platforms that are among the best solo software ever written. Inform 7 took Graham Nelson about three years from first design to the 2006 release, on top of a decade of Inform 6, and its manual is one of the glories of the field. TADS 3 took Mike Roberts the better part of eight years. Both are world-class work by people at the top of the craft, and both are a platform plus a language. Sharpee in 41 weeks is a platform, a language with its own compiler and intermediate representation, the platform's own grammar rewritten in that language, a native macOS IDE with a sealed toolchain, a website with a playground, and an eight-volume manual written in five days. The manual alone is the kind of thing that took the comparators years.

Complexity. This is not a CRUD application. A parser interactive-fiction platform is a world model with containment and scope, a natural-language grammar with disambiguation, a standard library of forty-odd actions each with validate, execute, and report phases and interceptors, an event system, save and restore, a multi-user server, and a language compiling to all of it. The 318 decision records are the size of the design surface. The three-architectures-in-thirteen-hours day is not someone typing fast; it is someone trying three answers to a hard design question in the time a human team spends scheduling the meeting about it.

Quality, where the mirror is honest in both directions. Industry delivered-defect density runs about one to seven per thousand lines. At that rate 353 KLOC of non-test code ships somewhere between 350 and 2,500 defects. Sharpee's record shows issue counts in the low hundreds and a verified catalogue of the ones that mattered in the dozens. Either the density is an order of magnitude better than industry, or discovery is limited by having one external user and one book-QA pass, and the honest answer is the second until proven otherwise. What the record does show is that the defects which got through are the kinds a human team ships under schedule pressure (a validator that returns early, a serializer that was never finished, a grader that defaults to pass) and that they were found and fixed in hours once a consumer ran the path. And no solo project in the domain has published an adversarially verified retrospective of its own record. Sixty-four corrected claims out of 168 is a quality number no comparator offers at all.

What does not hold up. A third of everything written was deleted. A human team would never have built the first server, the play-tree subsystem, or three action architectures in a day, because they could not afford to; the velocity bought exploration the before times forbade, and some of that exploration was waste. And the app missed the 5.0.0 publish by three days. The mirror's answer to that is that the before times would still be writing the grammar.

What it is

One person, with a harness, built in forty-one weeks what the same person could not build at all for a year and could not finish for two more. The eight weeks after the 5.0.0 publish say it was not a sprint: nine more publishes, three new model lines, no pause for any of them. The design did not change. What changed was that something could hold the design long enough to build it, and then build faster than any human-paced instrument could follow, which is what the first post in this series is about.

The retrospective's one-line summary of three and a half years is the one I'd keep: the design was never the bottleneck. What the project waited for was not a better idea. It was a collaborator that could hold the idea long enough to build it.

← All posts