ChatGPT, Claude, Grok September 3 Outage Wasn't Caused by Astra, but Exposed a Bigger AI Problem Few Discuss
Anthropic's Colossus 1 deal and other Memphis capacity used for Grok show how little of the infrastructure underneath is visible from the outside

Three major AI services went down within the same morning on 3 September, producing an unusual sequence of failures across ChatGPT, Claude and Grok. The timing quickly fuelled speculation that OpenAI's GPT-6 Astra rollout had overloaded its own systems and sent users towards rival models.
Anthropic first reported elevated errors around 6.23am Pacific Time, followed by Grok at about 6.30am PT. ChatGPT's main incident began around 7.43am PT, while Astra was announced later that day as a limited rollout.
Three Services, Three Causes
Anthropic described Claude's failure as an 'infrastructure issue' affecting Claude.ai, Claude Code, Claude Cowork and the Claude API. Service was restored at 16.16 UTC, or 9.16am Pacific Time, although Anthropic did not identify the component responsible.
Grok's outage lasted longer. Its status record showed the incident beginning at 13.30 UTC and ending at 17.05 UTC, while SpaceXAI later said the disruption followed 'an outage at our Memphis compute center' and apologised to its 'impacted compute partners'.
OpenAI gave a separate explanation for ChatGPT. A company spokesperson said a routing error beginning at about 7.43am PT made ChatGPT and Codex unavailable to some users, with a solution implemented at about 8.17am PT; the status incident was later marked resolved at 16.55 UTC.
The public record therefore points to three different failure descriptions rather than one confirmed technical cause. The timing also puts the start of Claude and Grok's incidents before OpenAI's Astra rollout.
The Memphis Connection Is Real, but Narrower Than It Looks
Memphis is where the infrastructure story becomes more complicated. In May, Anthropic announced an agreement to use all the compute capacity at SpaceX's Colossus 1 data centre, representing more than 300 megawatts and more than 220,000 NVIDIA GPUs.
Anthropic also said Claude ran across AWS Trainium, Google TPUs and NVIDIA GPUs, making Colossus 1 one part of a wider infrastructure footprint.
SpaceX's subsequent regulatory disclosures described a broader arrangement covering Colossus 1 and Colossus II, involving about 325,000 NVIDIA GPUs and a disclosed fee of $1.25 billion a month through May 2029.
The terms allowed either side to terminate after the initial three‑month period with 90 days' notice. But Colossus 1 is not shorthand for all of SpaceXAI's Memphis infrastructure.
Contemporaneous reporting distinguished Anthropic's Colossus 1 arrangement from other Memphis capacity used for Grok, so the public record does not establish that one facility failure affected both services.
SpaceXAI confirmed the Memphis outage for Grok and referred to affected compute partners, but did not identify Anthropic. Anthropic, meanwhile, never attributed Claude's outage to Memphis.
Why Astra Became the Suspect
The theory gained traction because OpenAI was preparing one of its biggest model announcements while three major AI services were experiencing problems. Posts on X speculated that Astra had either caused ChatGPT's failure or overwhelmed competing services after users moved away from OpenAI.
The chronology undercut that theory. Claude and Grok were already experiencing outages before ChatGPT's main incident, while Astra's 3 September release was initially limited rather than a global infrastructure switch, with broader access following later.
The Infrastructure Problem Underneath
The more consequential issue is not simply that AI companies share compute. It is that the dependency map beneath competing AI products remains difficult to see when something goes wrong.
Anthropic's infrastructure illustrates the point. Claude can run across AWS Trainium, Google TPUs and NVIDIA GPUs, while its SpaceX agreement gives it access to a substantial additional pool of capacity. SpaceX, meanwhile, operates infrastructure for Grok while also selling compute capacity to an outside AI company.
That does not make the companies a single technical system, and it does not show that Memphis caused Claude's outage. It shows how an AI market that appears competitive at the product level can have infrastructure relationships that are harder for users to map.
By 6 September, neither OpenAI nor Anthropic had published a full technical postmortem explaining the underlying failures. SpaceXAI had identified the Memphis outage behind Grok but had not publicly detailed the technical failure inside the facility.
© Copyright IBTimes 2026. All rights reserved.

























