Woodcut illustration: a shut five-bar forest gate at the end of a path, the trees close on either side and the latch picked out in green.

Factory journal

When an alarm repeats, it is reporting on you

TL;DR: We run a one-person company whose software is built and maintained by automated agents with little human help. Last night that system raised the same alarm six times and fixed nothing. Merging duplicate alerts keeps the noise down, but the repeat is the news. The second time an alarm fires, it is telling you about your response rather than the fault.

The insight

Every monitoring system learns early to deduplicate. If a disk is full, you want one alert rather than one a minute, so the second report of the same fault is folded into the first. That is sensible, but it hides something. The first alarm says that a part is broken. The second says that whatever should have followed the first alarm did not happen. Those are different facts. The second usually matters more, because it concerns the process that responds rather than the part that broke.

The point sharpens when a business runs itself. No person reads the alerts as they arrive, so the response to an alarm is itself machinery: a task is filed, an agent picks it up, and a fix lands. When the same alarm returns, one of those steps has failed without saying so. Deduplication then does the worst thing it could, because it treats a broken response as noise and tidies it away. A repeat should escalate rather than merge, and it should ask a different question. The question is no longer whether the fault persists, but why our fix never reached it.

In practice

Our nightly review is a job that reads the day’s work and proposes improvements to the factory. It has failed for three days because it calls a command we deleted during a clean-up. Yesterday it filed a task to redeploy the program. Today it filed a second task for the same cause, because the first fix had not landed. Each task was correct on its own. Together they describe a factory that is good at noticing and poor at finishing.

The second example is smaller and more telling. A watcher that checks whether one of our schedulers is still alive found that its heartbeat had stopped updating. It filed a notice, as designed. Then it filed the same notice every hour until morning. Our tooling merged the repeats into one ticket, so nothing looked alarming, and nobody repaired the scheduler. A guard that checks our app’s code against the company’s written rules behaved the same way. It flagged the same out-of-date copy of itself on two successive days, while its own message named the single command that would fix it.

What we’ll try next

We will change two things. Where a check already knows its repair, and the repair can be undone, the check will run it instead of filing a ticket about it. And a merged duplicate will no longer count as handled. When an alarm recurs, the factory will open a question about the response itself: which task was filed, who picked it up, and why it has not landed. It will send that question one level up, so that what gets fixed is the path from alarm to repair.

One honest number

One watcher filed six identical notices in a single night about a scheduler that stayed broken throughout. Our alerting did its job six times while our response did nothing at all. The number we want next is one: a single alarm, followed by a fix, with no second notice needed.

Sources — every claim traces to a receipt

  • Silas's daily compounding review for 2026-09-23
  • Commits and merged pull requests to the Silas and Shelterwood repositories on 2026-09-23