AssemblyAI incident

Outage on US Async Endpoint

Critical Resolved View vendor source →

AssemblyAI experienced a critical incident on September 16, 2026 affecting Asynchronous API, lasting 1h 14m. The incident has been resolved; the full update timeline is below.

Started
Sep 16, 2026, 08:22 PM UTC
Resolved
Sep 16, 2026, 09:37 PM UTC
Duration
1h 14m
Detected by Pingoru
Sep 16, 2026, 08:22 PM UTC

Affected components

Asynchronous API

Update timeline

  1. investigating Sep 16, 2026, 08:22 PM UTC

    We are currently investigating an issue that is affecting all customers using our Async API on our US endpoint. Users will be receiving a server error. We will update with more information as we learn more. This began at 8pm UTC.

  2. identified Sep 16, 2026, 08:24 PM UTC

    The issue has been identified and our team released fix. This is affecting approximately 50% of all async transcriptions. Success rate is beginning to improve.

  3. identified Sep 16, 2026, 08:57 PM UTC

    We are continuing to investigate these elevated errors. We will provide an update as soon as possible.

  4. identified Sep 16, 2026, 09:10 PM UTC

    We are seeing a reduction in the number of errors, but we are still working to fully resolve the issue.

  5. monitoring Sep 16, 2026, 09:12 PM UTC

    A fix has been released and we are continuing to monitor the situation. Our Async API service is recovering.

  6. resolved Sep 16, 2026, 09:37 PM UTC

    This incident has been resolved. Service on our Async API has returned to normal.

  7. postmortem Sep 21, 2026, 08:38 PM UTC

    **Summary** On September 16, 2026, between 20:04 and 21:08 UTC, our US Async API returned elevated server errors. At peak, roughly 50% of async transcription requests on the US endpoint failed, and a portion of the requests that did succeed took longer than usual to complete. Our EU endpoint and real-time/streaming services were not affected. **Root cause** The failure originated in AWS SQS, the managed queuing service our transcription pipeline uses to distribute work. Specifically, the Fair Queues feature began rejecting valid messages with an `InvalidParameterValue` error on a parameter our services had been sending successfully for months. There was no deploy or infrastructure change on our side that triggered this; AWS confirmed it as a service-side issue and rolled back the change on their end later that evening. Because the affected queues sit in the path between our API and our transcription workers, requests that could not be queued failed outright, and a portion of the work that did get through was delayed behind reduced throughput. We mitigated by disabling Fair Queues across our services and reverting to standard SQS queues, which restored normal operation roughly an hour after the first error. **What we're doing about it** * **Faster configuration changes.** Much of our service configuration lives in environment variables, which require task restarts to propagate. We are moving this to a global feature flag system so mitigations like this one take seconds rather than minutes. * **A single, unified kill switch.** Disabling the affected feature required toggling it service by service. We are consolidating this into one control. * **Expedited emergency deploys.** We are adding a reviewed break-glass path so urgent fixes can bypass non-essential CI steps. * **Multi-region failover.** We are prioritizing work to fail the US pipeline over to a second region when a regional provider dependency degrades, rather than relying on feature-level mitigations alone. We're sorry for the disruption. No customer audio or transcript data was lost, and any requests that failed during this window can be safely resubmitted.