Back to Insights
11 Jul 20264 min readAI Agents

Every agent could reach me, so all of them did

We wired a dozen autonomous processes to be able to message a human. Then the human's phone started buzzing every few minutes, and we learned — the hard way — a lesson that on-call engineers have known for twenty years.

The setup was reasonable on paper. We run a fleet of long-lived agent processes: workers that pick up tasks, a supervisor that watches them, a couple of desks that trade on their own schedule, health monitors, a queue reaper. Each one occasionally has something worth surfacing. So each one got a way to send a message to the owner. "Tell the human when something important happens" sounds like exactly what you want from an autonomous system.

What you actually get is a swarm.

The failure mode is structural, not sloppy

Nobody wrote a chatty agent. Every individual message was defensible: a worker finished a batch, the supervisor noticed a stalled line, a monitor saw a service blip and recover. In isolation, each one is the kind of thing you'd want to know.

The problem is that "important to me" is a per-process judgment, and there were a dozen processes each making it independently, with no shared view of what the human had already been told or how much noise had already gone out that hour. Two different daemons flagged the same underlying blip from two angles within a minute of each other. A worker announced it finished the very work whose start another had announced. None of them was wrong. The sum was unusable — and an unusable notification stream is worse than none, because the human stops reading it, and then the one message that actually mattered scrolls past unread.

This is alert fatigue. It's the oldest lesson in operations, and we walked straight into it from the agent side because the agents felt like colleagues, not like pagers. A colleague who texts you when something's up is helpful. Twelve colleagues who each text you when something's up, with no coordination, is a denial-of-service on your attention.

The rule that fixed it

We collapsed it to one line: no autonomous process messages the human directly. Not "message less." Not "raise the bar for what's important." A hard structural rule — the direct path from a daemon to a person simply doesn't exist anymore.

Everything a process wants to surface goes to one internal channel instead — a single hook that every agent posts to. From there, exactly one component is allowed to talk to the human, and its whole job is editorial: collect what came in, deduplicate, drop the routine, and emit at most one consolidated digest on a real change. The default is silence. The human hears from one voice, on a cadence a human can actually absorb, and that voice has the context to say "three things happened, here's the one that matters" instead of three separate buzzes.

The shift is from permission to architecture. We didn't ask each agent to exercise better judgment about interrupting a person — judgment doesn't compose across a dozen independent processes. We removed the capability and routed it through a single throttle that can see the whole picture. It's the same reason mature alerting pipelines have one Alertmanager in front of the pager and not a hundred cron jobs with the on-call's phone number.

What generalizes

If you're building anything with more than one autonomous process that can reach a human, the shape of the fix is worth stealing before you feel the pain:

  • One egress, not N. Every process emits to an internal bus. A single notifier owns the human channel. The number of things that can page a person should not grow with the number of things you run.
  • Deduplicate at the throttle, not at the source. A source can't know what the other sources already said. The aggregator can. That's the only layer with enough context to suppress the double-report.
  • Default to silence; earn the interruption. Make "say nothing" the resting state and a consolidated digest the exception. Reserve the direct, immediate ping for the genuinely rare case a human must act now — and route even that through the one notifier so it still can't storm.
  • Route the nudge, not the person. Most inter-agent "hey, look at this" belongs to another agent, not to you. Only what a human must decide should ever climb toward the human.

Autonomy is supposed to buy back your attention, not spend it faster. The moment every process in your system can tap a person on the shoulder, you haven't built an autonomous system — you've built a room full of interns with your phone number. Give them one manager who knows when to knock.

AI agentsmulti-agentalertingobservabilityautonomyon-call

Found this useful?