InfluxData incident
Query performance degradation in Azure Westeurope
InfluxData experienced a minor incident on August 25, 2026 affecting API Queries, lasting 1h 52m. The incident has been resolved; the full update timeline is below.
Affected components
Update timeline
- identified Aug 25, 2026, 01:49 PM UTC
The issue has been identified and a fix is being implemented.
- monitoring Aug 25, 2026, 02:29 PM UTC
A fix has been implemented and we are monitoring the results.
- resolved Aug 25, 2026, 03:41 PM UTC
This incident has been resolved.
- postmortem Sep 28, 2026, 01:00 PM UTC
# **Root Cause Analysis: Query Latency on August 25, 2026** Region: Azure West Europe Cluster: prod01-eu-west-1 ## **Summary and impact** On August 25, 2026, some customers in Azure West Europe experienced increased query latency and query timeouts. Additional query capacity helped reduce the initial queue, but full recovery required a manual restart of a stalled storage partition. Query latency returned to normal at 15:35 UTC. ## **Cause** A surge in query volume triggered scaling that quickly exhausted available resources. Shortly after the initial trigger, a storage partition did not complete startup successfully. Although additional capacity brought the query queue back to a nominal state, customer latency signals showed that recovery was incomplete. We discovered the stalled partition required manual intervention to complete initialization. One half of the partition remained online to serve queries and writes while its companion half was impaired. The underlying reason for the partition startup failure remains under investigation. ## **Recovery actions** The team allocated additional query resources, then manually inspected and restarted the impaired partition after identifying the incomplete recovery. The partition fully recovered at 15:32 UTC, and query latency returned to normal at 15:35 UTC. ## **Timeline** All times are UTC. ## **Future Mitigations** Improvements to the partition management workflow are in progress to strengthen resilience and automate recovery. Investigation into the startup failure continues. These improvements are ongoing work, rather than completed preventive fixes. We continue to strive for a healthy platform for all customers on our cloud infrastructure.