LiveKit incident

Investigating reports of failed SIP Transfer calls in US East

Minor Resolved View vendor source →

LiveKit experienced a minor incident on August 17, 2026 affecting US East - SIP, lasting 1h 11m. The incident has been resolved; the full update timeline is below.

Started
Aug 17, 2026, 06:10 PM UTC
Resolved
Aug 17, 2026, 07:22 PM UTC
Duration
1h 11m
Detected by Pingoru
Aug 17, 2026, 06:10 PM UTC

Affected components

US East - SIP

Update timeline

  1. investigating Aug 17, 2026, 06:10 PM UTC

    We're investigating an elevated rate of failures when transferring active SIP calls in the US East region. SIP calls themselves remain connected, and inbound and outbound calling are otherwise operating normally.

  2. investigating Aug 17, 2026, 06:22 PM UTC

    No transfer failures has been observed since 18:05 UTC and SIP transfers are currently completing normally. Impact was limited to two windows today between 15:48-16:08 UTC and 17:24-18:05 UTC. We're actively monitoring while we investigate the cause and put safeguards in place to prevent a recurrence.

  3. resolved Aug 17, 2026, 07:22 PM UTC

    We have identified the root cause of both spikes and have ensured safeguards to prevent a recurrence. We've been monitoring since 18:05 UTC and have seen no further transfer failures, and have confirmed that SIP transfers are operating normally. We will follow up with a detailed post-mortem.

  4. postmortem Sep 01, 2026, 11:44 AM UTC

    ## Summary On 17 August, between 17:24 and 18:05 UTC, a small percentage of SIP call transfers failed in our US East region during a routine configuration rollout. An issue in the automation that manages our SIP signaling servers prevented outgoing servers from being taken out of service safely, and those servers shut down while calls were still active on them. Transfers have completed normally since 18:05 UTC. Calls themselves stayed connected, and no other region was affected. ## Root Cause The configuration change restarts the servers that handle SIP signaling one at a time. Before a server is shut down, it is removed from service so that no new calls reach it, and it is then given some time to finish the calls it is already handling. In this case that removal did not complete expectedly and the outgoing servers kept receiving new calls for the entire drain duration, then shut down on schedule with calls still active on them. ## Timeline \(UTC\) * 16:52 - A routine configuration rollout begins in US East. * 17:24 - SIP transfer requests begin to fail for a subset of active calls. * 17:36 - Automated monitoring detects the elevated failure rate and our team begins investigating. * 17:40 - A further subset of active calls is affected. * 17:41 - Replacement capacity comes online. * 18:05 - Last of the impacted calls attempts a transfer and record a failure. ## Scope of Impact Only SIP call transfers \(`TransferSIPParticipant`\) in our US East region were affected, between 17:24 and 18:05 UTC. This represented 0.04% of all active calls in that window, and 1.1% of the calls that attempted a transfer. Customers who were impacted would have seen affected transfer requests return a 408; the underlying call stayed connected and only the transfer failed. Inbound and outbound calling were unaffected, as were calls that did not attempt a transfer, and no other region was affected. ## Mitigations and Follow-ups * We have deployed an alert for calls that end unexpectedly when a server shuts down. * We are changing our rollout process so that a server which cannot be removed from service safely halts the rollout. * We are preventing new calls from being routed to servers that are shutting down. * We are improving monitoring of the automation that manages SIP server rotation. * We are returning a more specific error when a transfer request cannot be delivered.