dsgnr.workKarthik S.UX + AI
← think/
an accountable AI system · systems story · Aug 6, 2026 · 8 min read

Every green light was honest

Four status checks said everything was fine. All four were telling the truth.

At about nine this morning I messaged the concierge on my Mac Mini and got nothing back. I could still reach the box over SSH, which is useful information on its own, because it means the machine is up and the private network is up and the problem lives somewhere in the last hop.

So I ran the checks, in the order a month of outages has taught me. The process manager said the gateway was running, and that it had never exited. The agent framework said the model provider was logged in. I asked the agent a question straight from the command line and it answered me immediately.

Four checks. Four greens. One dead bot.

Every one of those answers was true. That’s the part I’ve been sitting with all day.

the number/

The thing that broke it open isn’t a status indicator at all. Discord publishes a counter called the session start limit, which caps how many times a bot may identify itself in a day, and it only ticks down when an identify actually succeeds. Mine read 913 out of 1000.

Eighty-seven successful identifies in eighteen hours.

The bot wasn’t failing to reach Discord. It was reaching Discord, identifying, connecting, and then dying, over and over, all morning, while every status surface on the box reported that things were fine. One number, from a system I don’t own, killed every hypothesis I had about networks and DNS and revoked tokens in about four seconds.

Then it got worse. There are two log files. One has timestamps and one doesn’t, and I spent four rounds of reasoning on the one that doesn’t, which produced a confident and completely wrong story about a nine-hour outage. The timestamped file showed successful connections at 08:32, 08:46, 08:49, 08:52, 09:17, 09:20, 09:53. It hadn’t been down. It had been flapping, dropping and reconnecting every one to four minutes, and the window I happened to be staring at was just the longest gap of the morning.

I counted the failures by day. Fourteen days of zeros, and then every single day populated: 21, 37, 23, 21, 31, 14, 19, 98, 51, 62, 69.

I’d updated the agent framework from 0.18.2 to 0.19.0 on July 26.

Over the life of that log, 544 successful connections against 452 timeouts. Roughly a thousand attempts, about 45 percent of which failed, all of them after the update. It had been dropping and reconnecting around forty times a day for eleven days, in a file sitting on my own disk, on a machine whose every health check says running, and I noticed this morning only because the gaps finally got long enough to annoy me.

three times in one week/

What makes this a post rather than a bad morning is that it was the third time in a week.

On Tuesday I finally got a second model provider wired up, which is the fix for the single biggest structural weakness in the whole system: the concierge rides one provider, and I’ve written about the four separate ways that went wrong in July. One command added a fallback chain. The framework’s own command for listing fallbacks reported a healthy chain, primary in front, backup behind it.

It would have failed on every single request. The output token ceiling turns out to be a global setting rather than a per-provider one, and I’d deleted it during an unrelated change an hour earlier, correctly, because the primary provider has no ceiling that needs capping. So the backup silently inherited the framework’s default of 65,536 against a tier that rejects anything over 4,096. The insurance policy was a piece of paper. I only found out because I tested it, and testing it was awkward, because nothing in the tool lets you force a failover.

Then on Wednesday I went looking for the scheduled jobs, because my own documentation said there were three of them. The list command returned three. There are six. The list hides paused jobs entirely, not greyed out or marked, just absent, and three of the six were paused or superseded. A whole cost-reduction project I’d run weeks earlier had been sitting there working, invisible to the only command I ever used to look.

Two others I’d already found and half-forgotten. The command that prints the configuration under-reports it, showing two of eleven settings that a text editor and the web dashboard both show correctly. The command that lists credentials over-reports, merging what’s stored on disk with what it found lying around in the environment, so the file says one and the list says three. And a toolset showing as enabled with no credential behind it at all, which I wrote about in the last post, because it had been running on somebody else’s server the whole time.

what they have in common/

None of these tools lied to me. That’s what took me a week to see.

The process manager was answering “is this process running.” It was. The framework was answering “does this credential authenticate.” It did. The fallback list was answering “is this chain syntactically valid and does its auth resolve.” Yes to both. The cron list was answering “which jobs are scheduled to run,” and a paused job isn’t. Every one of those is a reasonable question with a correct answer.

The question I was asking was “is it working.”

Each surface answers something narrower than that, and the narrowing is invisible from the answer. Green doesn’t come with a scope note. Nothing in the word “running” tells you it means the process rather than the service, and nothing in “healthy” tells you the chain has never once carried traffic. You have to already know what the indicator declined to check, which means the indicator is most useful to the person who needs it least.

the part where I have to be honest about my own work/

For five years, the screen where a Webex Calling admin finds out what the AI did was my responsibility. I led design across those admin surfaces at Cisco, and in the last stretch that meant the AI features specifically.

The first post in this series called that a gap I couldn’t close from the design chair, which is true but vaguer than I can now be. What I didn’t expect was which part of the chair would turn out to be wrong. Not the permissions model, not the settings taxonomy, not the wording of the consent flows, all of which I’d have guessed and all of which have held up. The status line. The smallest, least-argued-about element on any admin screen I ever shipped, the one that gets designed in ten minutes because everyone already knows what it says.

This morning I sat in front of four of my own and got the wrong answer from every one, without a single one of them being wrong.

I don’t think the fix is more indicators. Three of these systems had plenty. The fix is that a status surface should say what it checked and when it last checked it, in the same breath as the verdict. “Running” should read as process running, last successful connection four minutes ago, which on this box would have been the whole diagnosis. A backup should not be allowed to say healthy until it has carried a request, and until then it should say never exercised, which is a different colour entirely. A list that omits rows should say how many it omitted. None of that is hard. All of it is the kind of thing that gets cut in a design review because the happy path reads cleaner without it.

The other half is that some questions can only be answered from outside. Every check I ran this morning was the system reporting on itself, and a process that stays alive by design when its connection dies will keep reporting alive, honestly, forever. The number that actually settled it came from Discord’s servers. When I design an accountability surface now, the first thing I want to know is which of its claims can be corroborated by something that isn’t the system making them.

what it saidwhat it meantwhat it couldn’t see
gateway runningthe process is alive87 reconnects in the previous 18 hours
provider logged inthe credential authenticatesnothing about reachability
fallback chain healthythe config parses and auth resolvesit had never carried one request
three jobs scheduledthree jobs are unpausedthree more, paused, doing the real work
Figure 01, four green lights, and the question each one was actually answering

I raised the connection timeout from 30 seconds to 120, which is the boring change the evidence pointed at: healthy connections on this box measure between 5.5 and 22.9 seconds against a hard 30-second cap, and every failure sits at exactly 30.000. A 22.9-second success is a failure waiting for a slightly worse morning. The framework’s own authors ship 180 seconds for a different chat platform, so they’d already decided somewhere that 30 was tight.

I left the connection health thresholds alone even though they’re my leading suspect for why the socket goes bad in the first place, because changing two things at once would have made the result unreadable.

And I wrote the fix down as applied, not verified. Nothing lets me force a marginal connection, so the only proof available is an absence of timeouts over the next few days, which I’ll go and count on Monday. Writing “fixed” in my own log today would have produced a perfectly honest green light of exactly the kind this post is about.