I recently gave a talk to computer science students about how my way of building software changed as AI went from suggesting lines to doing the work. This post is the written version. It follows one path, mine, in the order it happened.
The talk drew every stage as the same loop: plan, write code, open a pull request, review, merge. That is the software development lifecycle (SDLC), one change at a time. What changed from stage to stage was who owned each step. Before AI, I owned all of them.
Copy and paste
In spring 2023 I used ChatGPT. I pasted my code in, copied the answer back out and adjusted it by hand. The AI never touched the codebase. I was the only connection between the model and the code.
In fall 2023 the AI moved into the editor. I used Cursor mostly for autocomplete, accepting one suggestion at a time while I still planned, typed and reviewed everything myself.
In fall 2024 I moved to Cursor’s Composer, which could edit several files at once. I stopped typing the change and started accepting it. A few months later Andrej Karpathy gave that habit a name:
I “Accept All” always, I don’t read the diffs anymore.
Andrej Karpathy, coining “vibe coding”, February 2025
He also said it was fine for throwaway weekend projects. That caveat matters later in this post.
The agent runs the loop
For me, the real break came in fall 2025, when I started using Claude Code and Codex in the terminal. Both had been out since early 2025. A coding agent can read the codebase, run commands and see the result, in a loop. It could implement a change, run the tests, fix what failed and commit, all before I looked.
The feedback now came from the code instead of from me. The agent committed, and I reviewed and merged.
By winter 2026 I was running roughly three to five agents at once, each in its own terminal, switching between them. At work I kept a few long-lived checkouts of the code and reused them, one task after another.
Clean checkouts
Reusing those checkouts always felt a bit dirty. Their only real advantage was skipping setup. So in spring 2026 I started giving every task a clean checkout.
That meant every project had to set itself up from scratch:
- install its dependencies
- provide its configuration
- run its own database and tests without colliding with the other agents
- clean up after itself
This is the Twelve-Factor App, which I have followed for over a decade. Agents just made those principles pay off again.
Writing it down
I wanted the agents to do the work the way I do it. So whenever I noticed myself repeating something, I wrote it down as a skill: a written instruction the agent follows every time it does that kind of work. I had reusable commands from February and a skills repository from March.
The skills grew in the order I needed them. Planning, review, pull requests and test-driven development came first, then building and git. By July they covered the whole SDLC, from spec to ship.
Along the way the skills handed more decisions to the agent. The plan skill tells it to answer technical questions itself and ask me only about intent, priority and risk. The build skill runs unattended and stops when the scope changes or the behavior is unclear. A skill from September makes the agent set up a way to run the app and check its own work.
Letting go
In summer 2026 I stopped reading every change on my personal projects. Two things happened at once. The models got better, and I remember Opus 4.6 as the release that raised my confidence. I was also producing more code than I had time to review.
I wrote about that shift in My Workflow. An agent reviews the change, I read the pull request description, and I look closer when something feels off or the change is risky. At work every change still gets a human review.
Planning went the same way. I had planned with AI through dialogue for months, answering one question at a time. With newer models I gave less detail. The agent reads the code and researches first, and only asks me when it actually needs something from me. The plans were usually good.
Managing agents turned out to be a lot like managing people. Describe the result you want, not the steps, and trust the team to get there. It usually goes wrong when I start micromanaging.
Trust needs proof
Trust only works if the agent can show that its work is correct. For a bug that means reproducing it, fixing it, checking again and showing me the proof. My part shrinks to looking at the proof.
For that, agents need the tools I use myself: running the app, clicking through it, measuring and taking screenshots. Without proof, a loop just makes more work for me. My skill for this is adapted from Lauren Tan’s Loops You Can Trust.
Knowing what was built
Trust lets me hand over how something gets built. It does not tell me what was actually built. An agent can confidently build something different from what I think it built, and that gap is what slows me down now.
It happens every time I build something new and stop looking. While building my factory, agents spent days polishing parts before the smallest complete loop worked at all. The fix was to build that loop first, from start to finish, and only then add to it.
The same corrections kept coming up across projects:
- a workaround for the symptom, not a fix for the cause
- two pieces of code doing the same job
- the wrong abstraction
- a fallback nobody asked for
These compound, because the agent copies one bad example into every change that follows. In My Workflow I wrote that a model can produce a design that is coherent, specific and wrong. The checks keep passing on a design like that.
Ask it to explain
My answer was to ask the agent to explain what it built. Then I turned that into a skill, explain-diff, based on Geoffrey Litt’s version. A separate agent explains what the change is for, why it was built that way and what could break, and then it quizzes me before I ship.
The separate agent matters. The one that built the change would only repeat its own view of it. The quiz sets the pace:
A quiz is a speed regulator. Working with AI, it’s easy for the loop to run faster than the speed of human understanding.
Geoffrey Litt, Understanding is the new bottleneck, July 2026
The factory
Not knowing what my agents had built bothered me for a long time, so I converted my workflow into code. In mid-September I started dim-factory, a software factory that runs the SDLC from plan to ship.
It started as a skill reset. The factory began with no skills and only got the ones it needed, rewritten from what the old skills had taught me.
At its core is the record: one local database of everything my agents did and the commits that followed. Before an agent starts, it asks the record what was already done.
The goal is less of my time to ship the same thing. That means less re-explaining, fewer decisions handed back to me and less re-deriving what is already known.
Agents do the work that needs judgment, and the factory handles everything else. I sit at the gates that matter: anything hard to reverse, outward-facing or ambiguous waits for me.
The factory is still being built and is not ready for use yet. It has working parts, but I have not measured any time savings.
Gates over instructions
The clearest lesson from the record was about instructions. By mid-September I had typed “don’t add unnecessary comments” at least six times across 29 sessions, even though the rule was loaded in the agent’s instructions every time. The instructions did not hold. Gates, automatic checks that stop any change that breaks a rule, held every time. So the comment rule became a gate.
The rule matters more than it looks. A comment is where an agent explains away a workaround. I took the idea from Lauren Tan’s pstack, where the comment reviewer calls a long justification without a real exception “a confession”. With no comments allowed, there is nowhere to excuse a workaround, so the agent has to fix the cause. In my experience, the code gets better.
The general rule I follow now: anything a script can check becomes a gate, and anything that needs judgment goes to an agent with a clear brief.
Dim, not dark
A dark factory ships code that nobody read. The name comes from manufacturing, where some factories are so automated they run with the lights off. Addy Osmani describes the alternative:
A lit factory is the same pipeline with the lights left on where judgment lives.
Addy Osmani, Software Factories, Light and Dark, July 2026
Mine is dim. It starts light, and the record should show when I can switch off a light, one check at a time, based on evidence rather than a feeling. That part is designed but not built yet.
I would only turn the lights off for simple work. A task can run with nobody watching when it is quick to check and the check is hard to fool. Fixing a bug is much easier to hand over than building something new. Karpathy’s weekend-project caveat still holds: take the human out too early and you end up debugging code nobody understands.
Moving to the cloud
At the end of September I started moving my work from the terminal to the Claude app. It does not print all the tool output into the conversation, so I read what the agent says instead of everything it ran. Scrolling back through a long session is also much easier. Working in the app is less stressful than in the CLI.
The app also puts local and cloud agents in one place, and for me it is the only real option that does. Alongside cloud agents I use Remote Control, which lets me drive a session on my own machine from the app.
It is early, and I am still working out what to hand over. The first candidates are bug fixes an agent can verify on its own.
The clean checkouts turned out to help here too. Every project already sets itself up from scratch, so an agent in the cloud can start from a clean checkout the same way a local one does.
Walk it yourself
The talk ended on one point: you cannot copy someone’s workflow. You have to walk the path yourself to see how the pieces fit into your own.
A workflow converted into code might be the one thing you can copy. Copying a skill set gets you the instructions. A factory also runs the workflow. The slides are at talks.crisu.me/software-factory-2026 if you want the visual version of the path.