Building with a paper trail
Four living documents run this project, and their first readers are the agents.
Before the Mac Mini booted for the first time, its plan had already gone from v1.1 to v1.4 in two days of arguing with an AI about my own idea. That was July 1st and 2nd. The build started on the 3rd, and the plan is at v1.5.1 now. The gap between those version numbers is what this post is about: not the system the last post described, but the paper trail underneath it, which existed before the machine did and is the reason I can still tell you why anything on it is the way it is.
Four living documents run this project. A master plan states the current architecture in one place, versioned, with a changelog. A decisions log records every significant call and, more usefully, the why, so no future session re-litigates a settled question. A build journal gets a dated entry every working session. And a progress file exists to answer exactly one question, which is where do I resume.
| document | answers | rewritten? |
|---|---|---|
| master plan | what the architecture is now | yes, versioned (v1.5.1) |
| decisions log | why it was chosen | no, superseded entries stay |
| build journal | what happened that day | never |
| progress file | where do I resume | every session |
That sounds like project hygiene. The part that makes it more than hygiene is who the readers are: the first audience for these documents is the agents. Every AI session that touches this project starts by reading an orientation file that says what the project is, what the principles are, and how to behave in the folder. The personas on the Mini read these same documents as ground truth.
I wrote about a smaller version of this when I shipped a browser extension alone: a repo where the docs were the memory, because every AI session starts from zero and will cheerfully re-decide yesterday’s question. What changes here is the readership. Those docs were notes to myself that a model happened to read. These are read at boot by agents that take every word literally and have permissions, and writing for that reader is a different discipline than writing for a colleague who fills gaps with common sense. It has already changed how I write them.
rules that fell out/
A few working rules emerged that I didn’t plan for, and each one earned its place by something going slightly wrong first.
Chronological logs stay period-accurate. Early journal entries say “kai” and “Slack” even though that persona was renamed and Slack lost to Discord on July 5th. The old wording stays, with a note explaining the chronology, because the renaming is itself part of the story. Only the latest-state documents get kept current. Rewriting history to match the present is how you lose the trail.
Timestamps come from the shell, never from memory. An AI will confidently guess the time. It’s frequently wrong.
Placeholders don’t get to pretend to be decisions. The plan had several lines that read like settled choices (“secrets go in 1Password or env”) and were nothing of the sort. Those are now flagged as explicit design exercises, queued before their build steps. A plan should never let a placeholder wear the costume of a decision.
And exactly one owner per task, always. Two owners across separate async runtimes fork state and blur accountability. A consult returns a result to the owner; it doesn’t take the task over.
who catches what/
Documents this dense get errors in them, and the interesting question is which errors a machine finds and which it walks straight past.
Three from the first fortnight, all mine, all of a specific kind. The plan enabled SSH and Screen Sharing early in the hardening sequence; I pushed back, because opening doors before the environment is secure is backwards, and the plan got amended rather than followed. An auto-login dialog rejected a password the machine had just verified as correct, and the fix was noticing the dialog had pre-filled the username as “Kai” when the account is “kai”. A planning doc used the word “Curator” for two unrelated things, and the tell that it needed fixing was that even I misread my own document.
None of those are things I’m smarter about than the model. They’re things you catch by reading your own project with your own stakes in it. That, more than any prompt technique, seems to be the actual job now.
The other habit worth naming: verify against the source, especially when you feel sure. Before installing the agent framework I checked every assumption about how it behaves against its own docs. Three that felt obvious were wrong, and I’d have built around all three. That habit has since caught bigger things, and cost me more when I skipped it.
One I can preview, because it dates from this same stretch: the afternoon a prompt injection tried to get me to pipe a script into bash, hours before I installed Homebrew, whose official installer is exactly that shape. The difference is entirely where you got the command from. That one gets its own post.
For three days the progress file’s resume line said the same thing: install Claude Code on the Mini, then Hermes. It eventually moved. The system came alive, and everything since (the outages, the costs, the catches) went into the same four files, which is the only reason I can write any of it down now without guessing.