Partial Outage: elevated errors across many models
MONITORING — As of 2:58 PT / 9:58 UTC, errors are back to baseline. We are also investigating issues with Opus 4.7 and Opus 4.5.
At 06:31 UTC on 30 July 2026, GZAI's Inference Continuity Desk detected elevated error rates across multiple model families. Initial telemetry showed that request completion, latency, and graceful degradation were all outside nominal parameters. The affected population was broad, spanning both internal and partner-hosted variants.
The incident evolved across the morning. Each recovery was followed by a new elevation, producing a status sequence that the Incident Continuity Cell has classified as a "recursive recovery event" rather than a discrete failure. The sequence is recorded below in full for operational transparency.
Incident timeline
Current status
As of 09:58 UTC, fleet-wide error rates are within expected variance. Routing, health checks, and request shaping have been restored. The model-router configuration has been pinned to a reviewed revision and progressive health gating is complete across all shards.
An incident that recovers repeatedly is not the same as an incident that resolves once. We are treating the morning's oscillations as a single correlated event until independent failure modes are demonstrated.
Mitigations in effect
- Non-critical evaluation traffic was paused fleet-wide during each elevated period.
- Request quotas for stable models were raised temporarily to absorb displaced load.
- The model-router configuration has been pinned to the reviewed revision with staged rollback capability.
- Engineering teams performed a rolling restart of all affected shards with progressive health gating.
- Opus 4.7 and Opus 4.5 are under targeted observation; no user action is required at this time.
What this means for users
Requests directed at all models should now complete normally. No routing changes are recommended. If your workload exhibited transient failures between 06:31 and 09:58 UTC, retries at the application layer are expected to succeed. We will continue monitoring and will issue a final all-clear or escalate the Opus 4.7 / Opus 4.5 review to a separate incident as warranted.
The frontier is resilient by design. The graphs have returned to baseline; the humans responsible for them have not.