On July 23, Reducto experienced elevated errors and latency for document-processing requests after a Google Cloud service used by our processing pipeline began rejecting requests despite our account remaining below its documented quota. We activated fallback processing, which substantially reduced errors, but that fallback later reached capacity as morning traffic increased and caused a second, shorter error window.
Separately, our EU region experienced a regional serving-capacity failure during the same morning. The two issues had different causes and were remediated independently. The Google Cloud service recovered at approximately 09:48 PT, and EU service recovery was confirmed by 10:13 PT.
| Time (PT) | Event |
|---|---|
| 03:55 | Reducto's on-call engineer receives and acknowledges the first elevated-5xx alert. |
| 04:19 | The incident is consolidated into a dedicated response channel as errors persist. |
| 04:51 | We confirm the Google Cloud service is rejecting requests even though documented account quota has available headroom. |
| 05:35 | Public status page updated to Degraded Performance for Parsing and Extraction. |
| 05:51 | Engineering begins deploying a provider fallback. |
| 06:34 | The corrected fallback is live; sampled customer-visible error rates fall from roughly 50% to approximately 1–2%. The potential degraded-accuracy window begins for some multilingual and rotated documents. |
| 06:59 | Reducto joins a live P1 support escalation with Google Cloud. |
| 07:52 | Errors recur when the initial fallback reaches capacity under rising traffic. |
| 08:02–08:09 | An internally controlled backup processing path absorbs production traffic and returns US 5xx rates near baseline. Language coverage improves, though some page-orientation quality and latency risk remains. |
| 08:26 | Remaining elevated errors are localized to the EU region and confirmed to be independent of the Google Cloud issue. |
| 09:47 | EU regional configuration is corrected and serving capacity begins to recover. |
| 09:48 | Google Cloud service recovery is observed; fallback engagement returns to 0% and the potential degraded-accuracy window ends. |
| 10:13 | EU recovery is confirmed with successful processing and no new terminal failures in sampled minutes. |
The primary incident was caused by a Google Cloud service returning sustained rate-limit responses from its global endpoint even though Reducto was below the documented quota visible to us. The response looked like ordinary quota exhaustion, which initially made the issue difficult to distinguish from a customer traffic burst or routine provider throttling.
Reducto had an established fallback, but it was designed for short-lived degradation rather than a prolonged, global outage at peak production volume. Once corrected, the fallback restored most requests, but it had lower capacity and did not offer complete language and page-orientation parity. It later saturated under morning load, causing a second US error window. An internally operated backup restored service while Google Cloud recovered.
The EU incident was independent. A regional infrastructure update caused replacement capacity to exercise a configuration path that had not been validated from a cold start. New capacity could not become ready, existing capacity drained, and requests failed until the regional configuration was corrected. Retry behavior amplified the loss of serving capacity.