Case study 03 AI-assisted delivery

Running AI-assisted delivery like an operation

AI can write software quickly. It also forgets everything between conversations, sounds equally confident when it's wrong, and will happily change production. I built an operating system around it so speed didn't cost correctness.

My role
Designed the operating model and runs it every day
Since
Late June 2026
Assistants
A coding agent, a strategy chat assistant, a spreadsheet assistant
Release checks
For the October 5 release, on the exact tested version: 2,015 automated app tests passed (7 skipped) and all 74 database test suites passed

Four problems to solve

  • No memory. Every new AI session starts blank. Re-explaining the business each time is slow and drifts.
  • Several assistants, one truth. The strategy assistant can't see the code; the coding agent can't see the strategy chats.
  • Confident mistakes. An assistant reporting "done" isn't evidence that it's done.
  • Real consequences. People run the company on this system, so a bad change affects real work the same day.

The operating model

  • The repository is the memory. Decisions, definitions, open questions and a daily handoff are written down. If a chat and the repository disagree, the repository wins.
  • Every session starts the same way — confirm the state, read the rules and the handoff, and run a check that the last closeout still describes reality.
  • Every day ends with a closeout that records what changed and why, and compiles a self-contained briefing for the assistant that can't read the repository.
  • Every claim carries a label — verified fact, owner report, recommendation or unknown. That's why this site labels its claims too.
  • Every release passes gates: re-verify the exact reviewed version, run the full suites, apply and read back database changes, merge only that version, confirm the deployment, then check it live.
  • Risky commands are blocked automatically, and high-risk changes get a second review from a separate AI reviewer that sees only the change, not the conversation that produced it. That is a check on the work, not an independent security assessment.

What it looked like in practice

  • A review that said no. A separate AI reviewer, given only the change, rejected a launch-readiness change; the problems were fixed and re-verified before launch rather than argued away.
  • A release held at the gate. A reviewed change to refunds was held because my release condition — reproduce the disclosed limitation on the exact candidate — exposed a money-attribution defect. It was fixed before anything reached production.
  • Stopping on a mismatch. Before every production change, the starting state is checked against what was reviewed. When a handoff record turned out to be wrong, it was corrected in writing before anything else changed.
  • Approval in my own words. Production changes need my explicit approval each time — an assistant can't approve its own release.
  • Separate verification from acceptance. The assistant verifies a release live; I then test it myself, and the two are recorded separately.
  • One leadership page, kept current. An executive operations brief and an owner onboarding binder are reviewed at each daily closeout and published together, so leadership reads the same state the assistants work from.

What it made possible

A small company got a governed internal platform in weeks, not quarters, without losing track of why anything was built. The same method now runs outside work too — I use it for a personal AI coaching project.

Limits

The process has overhead, and it is only as good as the discipline to follow it. It doesn't replace an independent security review or real measurement of business outcomes; neither has been done yet, and both are planned.

Evidence and attribution

Where each claim on this page comes from. "Record" means dated release records; "My report" is my own account, not independently measured; "Plan" is not done yet. Private records stay private — they're available to discuss in an interview.

ClaimSourceType
Operating model in use since late JuneDated project recordsRecord
October 5 release checks: 2,015 app tests passed, 7 skipped; 74 of 74 database test suites passedRelease record for that version (the app's full local check and a fresh database rebuild)Record
A separate AI reviewer rejected a launch-readiness changeDated project recordsRecord
Refund release held at the gate until a defect was fixedDated decision recordsRecord
Leadership brief and owner binder published together at each daily closeoutPublication recordsRecord
Owner approval required for every production change; verification recorded separately from acceptanceRelease recordsRecord
The same method runs a personal coaching projectMy account; repository is privateMy report