Confluent Cloud Metrics API experienced elevated latency and error rates
Timeline · 1 update
- resolved Sep 17, 2026, 05:49 PM UTC
Confluent Cloud Metrics API experienced elevated latency and error rates between 9:55 AM PDT to 10:10 AM PDT.
Confluent had 54 outages in the last 2 years totaling 3303h 35m of downtime — averaging 2.2 incidents per month.
There were 54 Confluent outages since September 18, 2025 totaling 3303h 35m of downtime. Each is summarised below — incident details, duration, and resolution information.
Confluent Cloud Metrics API experienced elevated latency and error rates between 9:55 AM PDT to 10:10 AM PDT.
Starting on Tuesday, September 15th 2026 at 21:25 UTC, multiple Confluent Cloud services began entering a degraded state. We have identified the cause of the degradation and are actively mitigating the issue. Updates will be posted here in the next 30 minutes or when the issue is fully mitigated, whichever happens first.
The issue has been identified and a fix is being implemented.
The mitigation has been deployed and the degraded services are recovering as of 00:20 UTC on September 16th, 2026. We'll continue to monitor for any residual issues before resolving this incident in 1 hour (01:20 UTC).
This incident has been resolved.
We are investigating a service disruption affecting client connectivity in AWS us-east-1. Customers may see produce/consume failures
We are continuing to investigate this issue.
A possible root cause has been identified and the Confluent Cloud team is implementing a fix
A fix has been implemented and the team is monitoring recovery.
The issues is now resolved. Connectivity to AWS us-east-1 should be fully operational.
We are investigating elevated error rates and degraded performance for some services running in the Google Cloud (GCP) us-central1 region. This is related to an ongoing Google Cloud networking issue affecting a portion of the region. Customers in this region may experience intermittent connectivity issues, elevated latency, or errors. We are actively assessing the full scope of impact and will provide an update as more information becomes available.
The GCP issue affecting the us-central1 region has been mitigated and all Confluent Cloud services in the region have returned to normal operation. Provisioning capability, error rates and latency have recovered to expected levels. Customers should no longer experience impact. If you continue to see any issues, please contact Confluent Support.
We are investigating connectivity issues and elevated error rates affecting some Confluent Cloud clusters hosted in the Azure Southeast Asia region, beginning at approximately 03:13 UTC on 25 Aug 2026. Affected customers may experience degraded produce/consume availability or intermittent connectivity for clusters in this region. This is associated with an ongoing Microsoft Azure service disruption in the Southeast Asia region. We are actively engaged with Microsoft Azure.
We are seeing recovery of Confluent services in the affected Azure Southeast Asia region and are continuing to monitor the situation.
Confluent services in the Azure Southeast Asia region are now healthy and operating normally.
Between 21:15 – 22:30 UTC and 00:05 – 01:15 UTC on Aug 23–24, 2026, Confluent Cloud logging service was degraded. Customers may notice missing Connect and Flink logs in the Confluent Cloud UI, the Confluent CLI, and configured logging integration destinations during these windows. Other log types and the Connect / Flink workloads themselves were unaffected. The issue has been mitigated.
We are investigating an incident affecting services in the GCP us-west1 region. Customers with Kafka clusters in this region may experience elevated latency, request timeouts, or increased error rates. Our engineering teams are actively working to restore normal performance.
We are seeing recovery of Confluent services and are continuing to monitor the situation. The next update will be posted within 1 hour.
The GCP issues in us-west1 region have been mitigated. Confluent services are no longer seeing any impact based on our monitoring.
Flink statements are degraded in GCP us-central1 region. The impact started at 18:37 UTC on July 17. Currently, engineers are investigating the issue.
Engineers have identified and implemented a fix.
Currently, we are monitoring the results after rolling out the fix.
All the statements should be running now.
Confluent is aware of the ongoing AWS incident affecting eu-central-1. We have identified degraded performance impacting clusters in this region, and our engineers are actively working with AWS to mitigate the issue.
The issue impacting the AWS eu-central-1 region has now been resolved. All clusters are fully operational and healthy.
The issue has been identified and a fix is being implemented.
A fix has been implemented and we are monitoring the results
This incident has been resolved.
Kafka producers and consumers to some clusters in Azure spaincentral may not be able to produce and consume records from their topic partitions. An issue has been identified by Azure and is under investigation.
This incident has been resolved.
We are currently investigating this issue.
The issue has been identified and a fix is being implemented.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
We are currently investigating this issue.
Confluent Cloud Metrics API experienced increased error rates from 2026-06-25 15:15 UTC to 2026-06-25 15:55 UTC. The incident has been mitigated and the systems have recovered. We are continuing to monitor.
This incident has been resolved.
On June 23rd, between 7:00 to 7:16 UTC the Confluent Cloud Metrics API was unavailable. Confluent has identified the issue and restored functionality.
The team has identified the root cause of the issues affecting standard and basic Kafka clusters in the AWS us-east-2 region. We are actively working on a mitigation strategy and will provide further updates as they become available.
The team is now observing recovery. The team will continue to monitor the clusters to ensure full stability.
We are continuing to monitor for any further issues.
This incident has been resolved.
We are currently experiencing Intermittent Produce/Consume Unavailability in Azure Dedicated clusters. The problem started at 2026-06-16 13:50 UTC. We are currently investigating and will update as we know more.
We have found the root cause and have mitigated the problem as of 2026-06-16 16:09 UTC. We will monitor for any residual issues before resolving this incident in 1 hour.
No further issues have been observed. The incident is now resolved as of 2026-06-16 17:58 UTC.
No further issues have been observed. The incident is now resolved as of 2026-06-16 17:58 UTC.
We are investigating intermittent inter-region network connectivity issues affecting AWS us-west-2. Customers may experience delays or errors with Cluster Linking and Apache Flink processing for clusters in or communicating with this region. We are actively investigating the cause and working on mitigation.
We continue to investigate intermittent inter-region network connectivity issues affecting AWS us-west-2. Customers may experience delays or errors with Cluster Linking, Flink processing and logs for clusters in or communicating with this region. We are actively investigating the cause and working on mitigation.
We have deployed a mitigation for the intermittent inter-region connectivity issues affecting AWS us-west-2, which began at approximately 04:00 UTC. Customers with clusters in or communicating with this region may have experienced delays or errors with Cluster Linking, Flink processing, and accessing logs in the Confluent Cloud UI. We are monitoring to confirm full restoration.
This incident has been resolved.
We are currently experiencing increased error rates and latencies in the Azure westus2 region. This began at approximately 5:10 AM UTC 29th May. Our team is actively investigating it.
Confluent cloud services are experiencing a zonal outage in Azure West US2 linked to Azure network connectivity degradation. Team is continuing active investigation.
We are monitoring the health of impacted Confluent services in the Azure West US2 region. More details about the Multi Service degradation on the Azure status page: https://azure.status.microsoft/en-us/status
Confluent services in the Azure West US2 region are now healthy. We are continuing to monitor for residual effects as Azure continues to address their service outage. More details about the Multi Service degradation on the Azure status page: https://azure.status.microsoft/en-us/status
The Multi Service degradation on Azure has been mitigated. Confluent Cloud is operating normally.
Starting 20:26 UTC May 21, 2026, Confluent Cloud experienced connection failures in a set of PNI Gateways for Kafka Clusters in AWS us-east-1 and us-west-2. We have fully mitigated the issue as of 23:56 UTC May 21, 2026 and service is operating normally.
The incident has been resolved.
We are investigating data plane impact (produce/consume) resulting from elevated error rates and latencies from AWS us-east-1 resources in AZ4. EC2 instances hosted on this availability zone are impaired by loss of power during an ongoing thermal event.
We are monitoring the health of impacted brokers in the affected availability zone. More details about the thermal event on the AWS status page: https://health.aws.amazon.com/health/status
AWS has resolved the underlying infrastructure issue in us-east-1 (use1-az4). The majority of Confluent Cloud services have recovered and are operating normally. We are continuing to monitor for any lingering effects and performing cleanup of affected resources.
AWS has reported the underlying infrastructure issue in us-east-1 (use1-az4) as resolved. However, some Confluent Cloud customers in the affected availability zone continue to experience connectivity issues. Confluent is actively investigating and working to restore full connectivity for impacted customers.
The incident has been resolved and us-east-1 (use1-az4) is fully operational.
We are currently investigating this issue.
The issue has been identified and fix is being rolled out. ETA for rollout complete across impacted regions 2 hours.
Roll out of fix is progressing, expecting all clusters to complete in approximately 30 minutes.
A fix has been implemented and we are monitoring for next hour.
This incident has been resolved.
We are experiencing delays Tableflow external catalog sync in AWS. Our team has added a fix to resolve the situation and are monitoring the same.
This incident has been resolved.
AWS regional recovery is expected to be extended in both me-central-1 and me-south-1 regions. Customers requiring immediate restoration in these two regions are encouraged to review regional failover options.
Provisioning new workloads in AWS me-central-1 and me-south-1 regions is disabled. Migration steps to other AWS regions have been communicated to impacted customers.
Per AWS, the me-south-1 region is now completely unavailable. Customers will not be able to access clusters in the me-south-1 region. Customers who have workloads in me-central-1 are strongly advised to relocate them to an alternative region.
Resolved.
Confluent Cloud Metrics API was unavailable from 2026-03-05 03:23 UTC to 2026-03-05 03:58 UTC, the incident has been mitigated and the systems are functioning well. We are continuing to monitor and will provide an update on or before 2026-03-05 06:40 UTC.
Confluent Cloud Metrics API are now fully functional. The incident has been resolved and systems are functioning as normal.