- Detected by Pingoru
- Aug 03, 2026, 03:30 PM UTC
- Resolved
- Aug 03, 2026, 07:48 PM UTC
- Duration
- 4h 17m
Affected: Cloud Dashboard (cloud.livekit.io)
Timeline · 3 updates
-
investigating Aug 03, 2026, 03:30 PM UTC
We are currently investigating an issue where the sessions view does not populate in the Cloud Dashboard. Real-time connectivity is not affected and no data loss is observed.
-
investigating Aug 03, 2026, 04:51 PM UTC
We have traced the underlying issue to a backlog in our data pipeline, and we are continuing to investigate the root cause while working to clear the backlog. This data ingest delay results in the Sessions view returning no results for any time range ending at the current time, including all Quick Ranges options in the dashboard. As a temporary workaround, select "Custom range" and set the end time at least 30 minutes in the past to load your sessions. Real-time connectivity is not affected and no data has been lost.
-
resolved Aug 03, 2026, 07:48 PM UTC
We have fixed the data pipeline delay and all sessions should be loading now in the Cloud Dashboard.
Read the full incident report →
- Detected by Pingoru
- Jul 30, 2026, 10:55 PM UTC
- Resolved
- Jul 30, 2026, 10:55 PM UTC
- Duration
- —
Timeline · 2 updates
Read the full incident report →
- Detected by Pingoru
- Jul 30, 2026, 06:53 PM UTC
- Resolved
- Jul 30, 2026, 10:25 PM UTC
- Duration
- 3h 32m
Affected: Cloud Dashboard (cloud.livekit.io)
Timeline · 2 updates
-
investigating Jul 30, 2026, 06:53 PM UTC
We are currently investigating longer than usual dashboard loading times. No services appear to be impacted and we don't expect any data to be lost.
-
resolved Jul 30, 2026, 10:25 PM UTC
An internal one-off analysis job ran some expensive DB queries which slowed down load times of the dashboard's sessions page. The job was removed and load times are now back to baseline.
Read the full incident report →
- Detected by Pingoru
- Jul 28, 2026, 02:46 PM UTC
- Resolved
- Jul 28, 2026, 02:23 PM UTC
- Duration
- —
Timeline · 1 update
-
resolved Jul 28, 2026, 02:46 PM UTC
Between approximately 10:30 and 11:55 UTC on July 28, customers with telephony workloads in our London region experienced elevated error rates on SIP APIs. Small number of call transfers were affected. We mitigated by removing the affected infrastructure from service at 11:48 UTC, and error rates returned to normal by 11:55 UTC.
Read the full incident report →
- Detected by Pingoru
- Jul 27, 2026, 05:00 PM UTC
- Resolved
- Jul 27, 2026, 05:00 PM UTC
- Duration
- —
Timeline · 1 update
-
resolved Jul 27, 2026, 05:25 PM UTC
On July 27, our San Jose region experienced a minor networking incident beginning at 16:14 UTC. Users connected to this region may have experienced reconnects and increased session start times. Traffic was routed away from this region at 16:39 UTC and errors have since returned to baseline.
Read the full incident report →
- Detected by Pingoru
- Jul 23, 2026, 08:00 AM UTC
- Resolved
- Jul 23, 2026, 08:00 AM UTC
- Duration
- —
Timeline · 1 update
-
resolved Jul 24, 2026, 04:21 PM UTC
On July 23, between 07:50 and 08:07 UTC, media servers in our London region experienced a host-level network fault, causing them to fail health checks and restart. Users connected to these servers may have experienced dropped connections, or interrupted egress recordings. The affected servers were removed and replaced. The region has been operating normally since.
Read the full incident report →
- Detected by Pingoru
- Jul 14, 2026, 09:49 AM UTC
- Resolved
- Jul 14, 2026, 10:41 AM UTC
- Duration
- 52m
Affected: United Kingdom - Real Time Communication
Timeline · 2 updates
-
monitoring Jul 14, 2026, 09:49 AM UTC
Between 08:54 am and 09:06 am UTC, some requests to create new rooms in our London region failed with errors. Error rates returned to normal by 09:06 am UTC and a fix has been applied. We are continuing to monitor.
-
resolved Jul 14, 2026, 10:41 AM UTC
Error rates for room creation in our London region have remained normal since 09:06 am UTC and we are no longer seeing any issues.
Read the full incident report →
- Detected by Pingoru
- Jul 07, 2026, 10:30 AM UTC
- Resolved
- Jul 07, 2026, 10:30 AM UTC
- Duration
- —
Affected: Japan - Real Time Communication
Timeline · 2 updates
-
investigating Jul 06, 2026, 06:40 PM UTC
We are currently investigating increased connection failures affecting our Japan region beginning at 18:18 UTC.
-
resolved Jul 06, 2026, 06:54 PM UTC
Connection failures in our Japan region have returned to baseline as of 18:43 UTC.
Read the full incident report →
- Detected by Pingoru
- Jun 22, 2026, 01:05 PM UTC
- Resolved
- Jun 22, 2026, 02:40 PM UTC
- Duration
- 1h 34m
Affected: US Central - SIPUS Central - Real Time Communication
Timeline · 4 updates
Read the full incident report →
- Detected by Pingoru
- Jun 18, 2026, 05:49 PM UTC
- Resolved
- Jun 18, 2026, 07:10 PM UTC
- Duration
- 1h 21m
Affected: Cloud Dashboard (cloud.livekit.io)
Timeline · 3 updates
-
investigating Jun 18, 2026, 05:49 PM UTC
We are investigating reports of the Sessions and Agent Insights pages not being accessible on the LiveKit Cloud Dashboard. There is no data loss or ongoing session impact associated with this.
-
investigating Jun 18, 2026, 06:39 PM UTC
Session data (except Agent Insights) on the dashboard can still be accessed by changing the timeline on the Sessions page to "Past 7 days" or greater. We're still working to restore full access.
-
resolved Jun 18, 2026, 07:10 PM UTC
As of 18:45 UTC, all features of Cloud Dashboard are fully restored and functional.
Read the full incident report →
- Detected by Pingoru
- Jun 18, 2026, 05:09 PM UTC
- Resolved
- Jun 18, 2026, 05:40 PM UTC
- Duration
- 30m
Affected: US East - Real Time Communication
Timeline · 3 updates
-
investigating Jun 18, 2026, 05:09 PM UTC
We are currently investigating connection failure reports in US East.
-
monitoring Jun 18, 2026, 05:32 PM UTC
We have routed traffic away from US East region and connection failure errors are back to baseline. We are continuing to investigate and monitor SIP transfer errors.
-
resolved Jun 18, 2026, 05:40 PM UTC
Service is fully restored as of 17:18 UTC. Traffic is being successfully routed to other regions.
Read the full incident report →
- Detected by Pingoru
- Jun 16, 2026, 09:49 AM UTC
- Resolved
- Jun 16, 2026, 08:42 PM UTC
- Duration
- 10h 53m
Affected: India - Egress
Timeline · 3 updates
Read the full incident report →
- Detected by Pingoru
- Jun 11, 2026, 05:52 PM UTC
- Resolved
- Jun 11, 2026, 11:47 PM UTC
- Duration
- 5h 54m
Affected: Global Inference
Timeline · 3 updates
-
investigating Jun 11, 2026, 05:52 PM UTC
We are currently investigating intermittent 429 errors for Google Gemini models routed through LiveKit Inference.
-
identified Jun 11, 2026, 07:09 PM UTC
We have identified the cause as an upstream limit on our Google Gemini API account. Automatic failover to an alternate Gemini deployment is in place and has reduced the impact, but a subset of requests to Google Gemini models may still intermittently return errors when failover capacity is exceeded. We are actively engaged with Google support to restore full capacity and will share further updates as we have them. Other models and providers remain unaffected.
-
resolved Jun 11, 2026, 11:47 PM UTC
This incident has been resolved. The errors were caused by an upstream issue at our model provider (Google) that incorrectly triggered an account-level usage cap on our Gemini API access, returning rate-limit errors for a subset of Gemini requests routed through LiveKit Inference. Google identified the cause, rolled back the change on their side, and raised our account limits to prevent recurrence. Gemini requests are now serving normally and error rates have returned to baseline. Other models and providers were unaffected throughout.
Read the full incident report →
- Detected by Pingoru
- Jun 11, 2026, 03:30 PM UTC
- Resolved
- Jun 11, 2026, 03:30 PM UTC
- Duration
- —
Timeline · 1 update
-
resolved Jun 11, 2026, 03:50 PM UTC
Between 13:00 and 15:05 UTC, a subset of requests to certain Google Gemini preview models (gemini-3.1-flash-lite and gemini-3-flash-preview) routed through LiveKit Inference returned errors due to a project-level misconfiguration. These preview models did not yet have model failover configured; all other models and providers were unaffected. This was resolved by routing the traffic via an alternate provider and fixing the misconfiguration. As resolution measures, we are extending model-level failover to cover these models and adding increased monitoring for such failures in the future.
Read the full incident report →
- Detected by Pingoru
- Jun 11, 2026, 01:00 PM UTC
- Resolved
- Jun 11, 2026, 01:00 PM UTC
- Duration
- —
Timeline · 1 update
-
resolved Jun 17, 2026, 09:53 PM UTC
On 2026-06-11 at 13:00 UTC, a serialization bug was introduced via frontend components which rendered some services unavailable via the dashboard for some users. These services included SIP trunk and dispatch rule creation/management as well as several advanced features in Agent Builder. A fix was pushed on 2026-06-17 at 17:43 UTC which resolved the issue.
Read the full incident report →
- Detected by Pingoru
- Jun 05, 2026, 09:06 AM UTC
- Resolved
- Jun 05, 2026, 11:08 AM UTC
- Duration
- 2h 1m
Affected: India - SIPIndia - Ingress
Timeline · 3 updates
-
investigating Jun 05, 2026, 09:06 AM UTC
We're investigating elevated latency affecting Ingress and SIP services in our Hyderabad region.
-
investigating Jun 05, 2026, 10:05 AM UTC
As of 09:21 UTC, we've applied a mitigation by routing traffic away from the Hyderabad region. Failure rates are returning to normal levels. We're continuing to monitor the issue.
-
resolved Jun 05, 2026, 11:08 AM UTC
We identified a connectivity problem in the Hyderabad cluster. As mitigation, we've drained the cluster to route traffic away from it. Mumbai traffic is unaffected and serving normally. We'll undrain Hyderabad once we're confident it's functioning normally.
Read the full incident report →
- Detected by Pingoru
- Jun 03, 2026, 08:00 PM UTC
- Resolved
- Jun 03, 2026, 08:00 PM UTC
- Duration
- —
Timeline · 1 update
Read the full incident report →
- Detected by Pingoru
- Jun 02, 2026, 02:05 PM UTC
- Resolved
- Jun 02, 2026, 03:18 PM UTC
- Duration
- 1h 12m
Affected: US Central - SIP
Timeline · 5 updates
Read the full incident report →
- Detected by Pingoru
- May 28, 2026, 02:53 PM UTC
- Resolved
- May 28, 2026, 08:13 PM UTC
- Duration
- 5h 20m
Affected: Global Real Time CommunicationUS East - Real Time Communication
Timeline · 7 updates
Read the full incident report →
- Detected by Pingoru
- May 26, 2026, 07:53 PM UTC
- Resolved
- May 26, 2026, 09:32 PM UTC
- Duration
- 1h 38m
Affected: Cloud Dashboard (cloud.livekit.io)
Timeline · 4 updates
-
investigating May 26, 2026, 07:22 PM UTC
We are investigating reports of Agent Observability not accessible for some sessions.
-
identified May 26, 2026, 07:53 PM UTC
We have identified an issue causing partial data unavailability in the LiveKit Cloud Dashboard: For some users, Agent Observability is temporarily unavailable for some sessions. Some sessions in the 24-hour filter are showing as Active even though they have ended. A fix has been identified and we are backfilling the missing data.
-
monitoring May 26, 2026, 08:13 PM UTC
A fix has been implemented and we are seeing no more data outage for newly created sessions. We are monitoring the fix and are currently working on backfilling the missing data.
-
resolved May 26, 2026, 09:32 PM UTC
All missing data has now been backfilled and the incident is resolved.
Read the full incident report →
- Detected by Pingoru
- May 22, 2026, 08:26 PM UTC
- Resolved
- May 23, 2026, 04:36 PM UTC
- Duration
- 20h 10m
Timeline · 3 updates
-
investigating May 22, 2026, 08:26 PM UTC
We have observed network instability impacting cross-region traffic from Frankfurt beginning at 18:15 UTC. Traffic has been redirected from Frankfurt as of 19:55 UTC. Participants originally connecting via Frankfurt may see slightly increased latency until the region is restored.
-
monitoring May 23, 2026, 12:17 AM UTC
We have re-enabled our Frankfurt cluster and are monitoring for further disruptions.
-
resolved May 23, 2026, 04:36 PM UTC
We have been monitoring the clusters for the past 16 hours and have not found any further disruptions.
Read the full incident report →
- Detected by Pingoru
- May 15, 2026, 04:57 PM UTC
- Resolved
- May 16, 2026, 02:50 AM UTC
- Duration
- 9h 53m
Timeline · 4 updates
Read the full incident report →
- Detected by Pingoru
- May 07, 2026, 09:15 PM UTC
- Resolved
- May 07, 2026, 09:15 PM UTC
- Duration
- —
Timeline · 1 update
-
resolved May 08, 2026, 08:25 PM UTC
Between ~17:16 and 17:21 UTC on May 7, 2026, inter-region connectivity between our US East region and other regions was briefly degraded due to a network connectivity issue in our underlying cloud provider's data center. During this window, sessions that relied on cross-region track relay through US East may have experienced brief connection failures or reconnections for some participants. Our system recovered automatically once the cloud provider's data center network connectivity was restored.
Read the full incident report →
- Detected by Pingoru
- May 06, 2026, 07:46 PM UTC
- Resolved
- May 06, 2026, 11:04 PM UTC
- Duration
- 3h 17m
Affected: Global SIP
Timeline · 5 updates
Read the full incident report →
- Detected by Pingoru
- May 06, 2026, 04:01 AM UTC
- Resolved
- May 06, 2026, 09:41 AM UTC
- Duration
- 5h 39m
Affected: US East - Cloud Agents
Timeline · 9 updates
Read the full incident report →