Skip to main content

Back on Track

Jul 24, 2026 · 10 min read

After stepping away from Acolyte, I rebuilt the development loop around daily dogfooding: fix completion and transcript integrity, delete scaffolding, and ship only what real use proves. The project is moving again without pretending it is finished.

0:00

I had burned out building Acolyte, and the project had grown beyond what I could keep in my head. The month-long pause I described in Coder’s Block stretched into roughly two months away. By early July, I decided to return.

The problem was not a single bad subsystem. The code could pass its checks while behavior decayed outside my view because I kept shipping features without using the tool on itself. By the time I noticed, the project felt too large to hold.

I needed a different way to try again. Time does not fix a project that feels too large. A method can. Over the last few weeks I dogfooded the tool, followed the failures into the code, removed machinery that no longer earned its keep, and kept going until the next step was clear again.

The models had improved while I was away. When I returned, they could handle more of the reasoning and implementation than they had before. That made a systematic recovery possible, but it also made plausible answers easier to trust too quickly.

Start with the bar

I asked Claude to gather every defect it could find, and that list became the plan. The work was turning that inventory into a bar: the behavior Acolyte had to get right before it was worth trusting again. It had to do what I expected, retain enough context to stay coherent, and preserve transcript content and intra-turn order.

That priority order changed the work. Behavior came first, then transcript integrity; continuity and memory could wait until those two held. Features did not outrank failures found in daily use. The plan’s job was not to describe the whole project. It was to make the next bounded win obvious and stop me from relitigating the priority every session.

I tracked the work in one markdown plan file, not a backlog of GitHub issues. Issue tracking earns its place once work is coordinated across contributors; for a solo recovery, it sat between me and the next fix.

The plan was a living document, not a log. When I found a bug or thought of an improvement, it went into the file, captured against the priority instead of pulling me off the current fix. What shipped was reconciled back into the plan, while git history stayed the archive of what had already happened.

Repository rules kept the invariants and commands in view while I worked, so each session started from the same bar instead of re-deciding it. I was not trying to finish Acolyte. I was trying to make it trustworthy enough to keep using.

Use the product

The first real task was using Acolyte to work on Acolyte. That sounds obvious for a coding agent. It was also the thing I had avoided for long enough to burn out.

Dogfooding changed the feedback loop. A test can tell me that a function returns the right value. A real session can tell me that the agent lost the answer, rendered the tool output in the wrong place, or stopped before doing the work. Those failures are not interchangeable. The second kind is what determines whether the product is useful.

Claude did the diagnosing, working from the trace. Acolyte records every session’s lifecycle and model calls, so Claude could follow a strange session, separate a model decision from a harness failure, and fix the boundary where it actually broke. The work settled into a rhythm: reproduce one behavior, follow it through the trace, fix it at the source, add a regression test, move on. That left me the judgment: which failures were real, what to fix first, whether a fix actually held. With the agent moving fast, deciding those was where the real work sat.

The bigger unlock was letting Claude operate the product itself. Before, Claude could reason about the renderer but never see what the terminal drew, so every visual check came back to me: run it, look, report what broke. The /dogfood skill removes that step. acolyte run could already exercise Acolyte headlessly, but it never touches the chat, where the rendering failures lived. Driving the live chat through tmux does, so Claude runs a real session and checks the transcript itself.

The order mattered. Completion came before transcript rendering because a session that cannot finish honestly is not useful. Transcript integrity came before auth, filesystem layout, documentation, or benchmarks because those improvements only matter when the core loop can be trusted.

Fix completion first

The first fixes were about the agent finishing a turn honestly. For most of the project’s life, completion was something the model had to announce, and the host, the harness around the model, propped that up with forcing and overlapping gates. When a model would not announce cleanly, the host tried to manufacture the ending. Deciding whether the work is done is the model’s call, and the host was the wrong place to make it.

The rework moved completion onto the model’s native end_turn. A step with no tool calls, carrying text, is the final response. The host stopped forcing completion and instead classifies that terminal step by its finish reason and answer text. A normal finish with text is accepted. A blank or truncated step is reopened once, each reason getting a single retry, and a cut-off answer is stitched back onto its earlier fragment. A content filter or a provider error ends the turn, with the host writing the user-facing message rather than letting a broken step pose as the answer. A model that returns a real answer never touches any of this.

The system prompt changed to match. Acolyte’s prompt now states the stance plainly: nothing forces its hand, and it stops when it knows the work is done. The old coaching toward a finish line, and the nudge to review its work before ending, came out. Keeping them would have been a leash the host had already dropped.

That was a large deletion as much as a lifecycle change. The forcing gates went, the duplicate completion gate went, and the host-side heuristics that guessed at completion went with them. Tool errors no longer block a turn from ending. The system became easier to reason about because the host stopped trying to manufacture certainty. The model decides when it is done. The host only makes sure a missing or cut-off answer is visible instead of pretending the task succeeded.

Repair the transcript

The other half of the recovery was the terminal itself. The renderer had been producing failures that are easy to miss in a unit test: duplicated scrollback, repeated-resize corruption, and output arriving in the wrong order.

The fix was not another rendering special case. It was a headless terminal test harness, tool output moved onto the ordered stream, clipping at the render layer, and regression coverage for transcript integrity. The system now preserves the ordered transcript data and decides how much to show at paint time. That is a safer boundary than truncating the data before the trace or the model can see it.

That work is now closed. Tool output moved onto the ordered stream, so within a turn the transcript records what happened in the order it happened. The presentation pipeline was then rebuilt to run in one direction: publish the semantic transcript, lay it out, then render it, with no path back from rendering into the data. A visual change can no longer reorder or corrupt what the transcript holds. The duplicated scrollback, resize corruption, and out-of-order output that opened this section are gone.

Delete the machinery

The same work produced several deletions, and they shared a rule: if real use did not justify a piece of machinery, it came out.

Two were whole subsystems. I removed the multi-file code-edit path and the orphaned diff-rendering subsystem left behind with it. Both added complexity the actual work never exercised, and both raised the cost of every change made near them.

The rest was machinery that steered the model from the host: a reminder subsystem that nudged it mid-turn, a breaker that gave up on a turn after a run of tool failures, and keyword-triggered skill activation that guessed which skill the model needed. Each substituted a host rule for the model’s judgment. A failure is information the model can act on, not a reason for the host to abort. A skill the model has not reached for should not be forced on it by a keyword match. Removing them made the model’s decisions visible instead of burying them under another policy layer.

The point was not fewer features. It was fewer decisions the host made for the model.

Put it back in use

Once the core path was reliable enough for bounded use, Acolyte v0.22.0 shipped on July 20, with patch releases following. I updated the website alongside it by syncing the canonical docs, refining the terminal demo, and aligning the benchmark presentation.

Those changes only mattered because the core path was becoming usable again. A cleaner auth command does not compensate for a lost transcript. A better benchmark does not prove that the agent can finish a task. The surrounding product surface only matters once the real product is reliable.

The operating bar is simple now: use the product, run the relevant checks, review the evidence, and do not merge behavior changes without dogfooding them. A change is trustworthy when the result is reproducible and the remaining uncertainty is in the open.

Back on track does not mean finished. It means the failures that remain surface in daily use, each with a reproduction and a verification gate, instead of decaying out of view. The next milestone is not 1.0. It is Acolyte being solid enough that someone else can try it and tell me where it breaks. The project is moving again because the next fix is visible in the product, not only in its commit history.

Share

Read next

Paint Last

Acolyte's chat packed three jobs into one place: what to say, how to lay it out, how to color it. So I could not redesign the look without risking agent behavior or the transcript. The fix is a one-way pipeline where paint comes last and cannot reach back.