elmerdata.ai blog

My blog

Who Checks The Outage?

On September 3, ChatGPT, Claude and Grok failed within one morning. No shared cause was proven, and no one has the authority to look for one.


no_ones_problem Someone has to be responsible for the mess. Irresponsible littering, Salesbury, Lancashire, England, 2024. Photograph by philandju / Geograph Britain and Ireland. CC BY-SA 2.0.


On the morning of September 3, three of the four frontier AI services failed within about an hour of one another. Grok went down at 6:30 a.m. Pacific and stayed down for three and a half hours. ChatGPT and Codex began erroring at 7:43 and recovered by 8:17. Claude ran a partial outage of three hours and six minutes, with Opus 4.8 and Opus 5 hit hardest, and restored service at 16:16 UTC. Gemini wobbled but issued no outage notice. The evidence does not establish that one Azure region, one GPU cluster or one shared network failure brought all three down together. It is messier than that, and more interesting.

Each company explained itself in a few words. OpenAI cited "a routing error." Anthropic cited "an infrastructure issue" and said nothing further. SpaceXAI apologized for "an outage at our Memphis compute center" and, in the same statement, apologized to unnamed "compute partners." Cloudflare denied involvement outright, and the status pages at Azure, AWS and Google Cloud showed nothing. Microsoft did not comment. The Azure theory that circulated that day rested on third-party crowd reports of ingress failures in the East US region, not on anything Microsoft said.

The fault line on the public record

One fact the coverage mentioned in passing belongs at the center. Since May 6, Anthropic has leased the entire compute capacity of Colossus 1, the SpaceXAI facility in Memphis, at more than 300 MW across roughly 220,000 Nvidia GPUs. SpaceXAI says its Memphis compute center failed. SpaceXAI says compute partners were affected. Anthropic says it had an infrastructure issue. No one has connected those three statements in public, and no one is obliged to. OpenAI's routing error looks independent of Memphis, so the honest reading of September 3 is not "three unrelated failures" and not "one hidden cause." It is two houses that may share a foundation, one house that apparently does not, and one owner declining to say which.

That is the concentration argument in miniature. Modern AI services rest on overlapping cloud regions, compute partners, GPU supply and network backbones, so a service that looks independent at the application layer can share much of its foundation with a competitor underneath. The lease makes the abstraction concrete: two rivals, one building, one bad morning.

What every other critical sector already has

The striking thing about September 3 is not that the companies were vague. It is that vagueness was their choice to make. In every other sector where correlated failure carries public consequences, some authority owns the question of what happened.

After the August 14, 2003 blackout, a joint US-Canada task force published a root-cause report in April 2004 naming four causes, from untrimmed trees to a failed alarm system at FirstEnergy, and its first recommendation was to make reliability standards mandatory. Congress did that in the Energy Policy Act of 2005, and NERC standards became enforceable in June 2007. Telecom carriers have filed outage reports with the FCC's Network Outage Reporting System since 2004, on thresholds measured in user-minutes, whether or not anyone is asking. Exchanges, clearing agencies and the largest trading platforms fall under Regulation SCI, which since 2015 has required them to notify the SEC of systems disruptions and to review the cause. In Europe, the Digital Operational Resilience Act, in force since January 2025, extends that logic one layer down: it reaches the ICT third parties on which banks depend, which is precisely the compute-partner layer that failed on September 3.

None of that reaches an AI lab. Frontier model providers are not designated critical infrastructure in the United States. They carry no mandatory incident-reporting duty, no root-cause obligation and no regulator with standing to ask for either. When the grid fails, someone with subpoena power finds out why. When a broker's systems fail, the SEC gets a notice. When three AI services fail in one morning, the public gets three status-page sentences and a post promising corrective action.

The gap is the absence of an investigable event

The governance point is easy to state too weakly. The problem is not that someone should have investigated September 3 and did not. It is that no such thing as an investigable event exists yet for this layer of the economy. The firms alone decide what counts as a cause, how much of it to disclose and when the matter is closed. Anthropic's decision to say "infrastructure issue" and stop, as the tenant of a facility its landlord says failed, is not a scandal under any current rule. It is the rule, because there is no other.

Several earlier posts here have circled that seam from different directions. The SEC's silence on agentic trading leaves deployment responsibility unassigned among broker, developer and customer. Governance arrived too late for agentic AI because the frameworks were built for a previous generation of systems. When the Firm Becomes the Monoculture showed a single company reproducing correlated failure inside its own walls, with outages as the evidence. September 3 adds the missing piece: responsibility is unassigned even for the question of what happened. A sector can run a long time without rules for how to behave. It cannot earn trust without rules for how to find out.

The enterprise lesson is already being drawn, correctly, as a resilience problem: run more than one model, run more than one provider, know who your provider's providers are. The advice is sound and beside the point, because a customer cannot diversify away from a dependency it is not allowed to see. Anthropic's customers did not know on September 3 whether their outage and Grok's were the same outage, and they still do not.

Three houses shook on the same morning. No one found one earthquake beneath all three, but no one was sent to look, and at least two of the houses appear to stand on the same ground. The lesson of September 3 is not that the industry shares fault lines. It is that no surveyor has the authority to map them.


Further Reading

From this blog

Sources


AI Assistance Statement ▾

This blog publishes at a near daily rate, and that pace is possible because AI tools do a substantial share of the work between the idea and the published text. Preparation of this entry included assistance from Anthropic's Claude and from OpenAI's ChatGPT (GPT-5 series reasoning models). I use them to research a topic and gather primary sources, to organize ideas and propose structure, to draft and revise prose, to check factual claims against the cited sources before publication, and to score drafts against a set of house style rules. Longer pieces are often developed across several sessions. A written handover carries the argument, sources, and open questions from one session to the next, and the same tools help prepare those handovers. The tools also help identify candidate images and confirm that selected images appear to be released for reuse, for example through public domain or Creative Commons licensing.

The process is also an AI experiment in its own right. There is a live argument about what AI-assisted writing does to originality, and whether the result is thought or slop; the August 2026 dispute over a Wall Street Journal op-ed that its author acknowledged drafting with AI, and the Journal's subsequent defense of the practice (WSJ is behind a paywall, but for a public summary: click here, and here), is one newsworthy example. I would rather run the experiment openly than pretend it is not happening. This blog is one sustained attempt to find out whether a person with an argument, working with these tools every day, produces writing that is still recognizably that person's, and I disclose the method so readers can judge the result.

The judgment is mine. I choose the topic, decide the argument, supply the personal and professional experience the pieces draw on, read and edit every draft, verify the sources and image licensing, and take full responsibility for the final published content. Where a post contains my own recollections, the AI did not invent them.

Statement revised September 2026.


#AIData #Observations