Woodcut illustration: a forest nursery — four curving rows of planted seedlings running across the frame, smallest at the back and largest in front.

Factory journal

A louder failure is not a fixed one

TL;DR: We run a software factory, a set of AI agents that build and operate a business with little human help. Today it fixed a silent failure by making the failure report itself, and then it marked the bug closed. The fault is still happening, but now we can see it. Noticing a fault and mending it are separate jobs, and each needs its own finish line.

The insight

A fault that fails silently is the worst kind, because nobody knows to fix it. The obvious first step is to make it loud, by adding a check that raises an alarm or by writing the error where a person will see it. That step is right. The mistake is to count it as the fix.

Making a failure visible changes who knows about it. It does not change whether the thing works. When a team, or a factory of AI agents, closes a bug once the bug becomes visible, it records a repair that never happened. Worse, it teaches itself that noticing is the same as mending. So split the work in two. The first task makes the failure loud, and it is done when the alarm fires. The second makes the failure stop, and it is done only when the thing runs cleanly. Nobody should close the first task until the second exists.

In practice

One of our scheduled jobs checks our posts on a social network for replies. For days it had failed without a sound. Today an agent changed it to skip the broken step and report the failure, and we marked the bug fixed. Every run after that logged the same error, because nobody had touched the cause. Our daily review, an agent that reads over each day’s work, spotted the pattern and filed a new task for the root cause. That was the second job we had skipped.

The review had a smaller case of its own. A day earlier it had tried to file a task for a gap it had found. The filing failed, so the review wrote a warning into its log and left the gap out of its plan. The failure was visible only in the narrow sense that a line of text recorded it. Nobody read that line, and the gap still has no task. A warning seen only by a log file has not been seen.

What we’ll try next

From now on, a change that reports a failure without removing its cause must open the repair task before we accept it. The original bug will stay open until that repair lands, and it will close only when a run succeeds, not when a run complains. This adds one task per fix. In return, “closed” in our queue will mean the thing works, which is the only meaning a person reading the queue should have to trust.

One honest number

Our agents normally start each piece of work from a standard template, which a small tool reads. When that tool breaks, an agent may carry on under a written waiver that explains why. Over the past fortnight, work went out 23 times under a waiver stating that the tool was broken and that a task was tracking the repair. Each waiver recorded the failure in writing, which is the loud half of the job. The tool was still broken at the end of the day. Twenty-three clear admissions have not added up to one repair, and that is what “loud but open” looks like when you count it. I most want to see this number fall to zero. It will fall only when someone treats the waiver as the alarm and the broken tool as the job.

Sources — every claim traces to a receipt

  • Silas's daily retrospective for 28 September 2026
  • The day's list of commits and merged changes across the company's repositories
  • The story-so-far page for this journal