- Detected by Pingoru
- Sep 15, 2026, 04:08 AM UTC
- Resolved
- Sep 15, 2026, 05:34 AM UTC
- Duration
- 1h 25m
Affected: Deployments
Timeline · 4 updates
-
investigating Sep 15, 2026, 04:08 AM UTC
We are investigating reports of Depot builds failing for some customers. Affected customers can deploy successfully using the --depot=false or --buildkit arguments to flyctl.
-
identified Sep 15, 2026, 04:45 AM UTC
We have identified an issue with deploying via Depot for users connecting through our SYD and JNB regions. Affected customers in these regions can deploy successfully using the --depot=false or --buildkit arguments to flyctl.
-
monitoring Sep 15, 2026, 04:51 AM UTC
A fix has been implemented and we are monitoring builds. Standard flyctl builds should be working again for customers in all regions.
-
resolved Sep 15, 2026, 05:34 AM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Sep 12, 2026, 09:22 PM UTC
- Resolved
- Sep 12, 2026, 10:12 PM UTC
- Duration
- 49m
Affected: LAX - Los Angeles, California (US)SJC - San Jose, California (US)
Timeline · 3 updates
-
investigating Sep 12, 2026, 09:22 PM UTC
We are investigating upstream network issues from US West Coast (SJC, LAX). Apps hosted in US West regions may experience higher latency or packet loss, and requests from clients physically located in US West may experience higher latency.
-
monitoring Sep 12, 2026, 09:57 PM UTC
Private networking between Fly Machines is resolved, and most outbound connections are healthy. We're continuing to monitor the network, and some issues will still be expected from clients physically located in US West until upstream transit issues are resolved.
-
resolved Sep 12, 2026, 10:12 PM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Sep 02, 2026, 11:03 PM UTC
- Resolved
- Sep 03, 2026, 12:41 AM UTC
- Duration
- 1h 37m
Affected: Sprites
Timeline · 3 updates
-
investigating Sep 02, 2026, 11:03 PM UTC
We're aware of a problem affecting a subset of Sprites users. We are investigating the source of the issue.
-
monitoring Sep 02, 2026, 11:33 PM UTC
Error rates have decreased. We are continuing to monitor the API health.
-
resolved Sep 03, 2026, 12:41 AM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Sep 02, 2026, 02:18 PM UTC
- Resolved
- Sep 02, 2026, 03:57 PM UTC
- Duration
- 1h 39m
Affected: LAX - Los Angeles, California (US)SJC - San Jose, California (US)
Timeline · 5 updates
-
identified Sep 02, 2026, 02:18 PM UTC
We have observed an upstream network issue in LAX. Connections to some destinations may see elevated latency and packet loss. We're working with our upstream to resolve this issue.
-
identified Sep 02, 2026, 02:48 PM UTC
We have put in some temporary mitigations along with our providers. However since the root cause of this issue lies within a bigger upstream transit provider, you may continue to see some elevated latency and connection issues in/around affected regions. We're still working closely with them to resolve the root cause.
-
identified Sep 02, 2026, 02:56 PM UTC
We're updating the affected region list to also include SJC since this seems to be a wider upstream issue in US West Coast.
-
monitoring Sep 02, 2026, 03:33 PM UTC
A fix has been implemented and we are monitoring the results.
-
resolved Sep 02, 2026, 03:57 PM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Sep 02, 2026, 07:21 AM UTC
- Resolved
- Sep 02, 2026, 07:44 AM UTC
- Duration
- 22m
Affected: Machines API
Timeline · 3 updates
-
investigating Sep 02, 2026, 07:21 AM UTC
We are investigating an issue with the background job runner for our API. Actions that require a background job, such as creating apps, assigning IP addresses, or creating/renewing certificates, may fail at this time.
-
monitoring Sep 02, 2026, 07:31 AM UTC
A fix has been implemented and we are monitoring the results.
-
resolved Sep 02, 2026, 07:44 AM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 31, 2026, 03:42 PM UTC
- Resolved
- Aug 31, 2026, 03:30 PM UTC
- Duration
- —
Timeline · 1 update
-
resolved Aug 31, 2026, 03:42 PM UTC
A configuration update caused temporary failures for incoming HTTP/2 traffic for Fly Machines located on a subset of hosts for a few minutes. This incident has since been resolved. Managed Postgres depends on HTTP/2 and some control plane ops may have been affected as well. However, downstream Postgres connections were unlikely to have been affected by this since they do not use the HTTP2 handler.
Read the full incident report →
- Detected by Pingoru
- Aug 31, 2026, 06:26 AM UTC
- Resolved
- Aug 31, 2026, 08:10 AM UTC
- Duration
- 1h 44m
Affected: ORD - Chicago, Illinois (US)
Timeline · 3 updates
-
investigating Aug 31, 2026, 06:26 AM UTC
Due to an upstream provider, we are seeing ~50% packet loss on a subset of hosts in ORD. Some MPG clusters in ORD are slow to replicate as a result.
-
monitoring Aug 31, 2026, 07:39 AM UTC
Packet loss in ORD is improving and impacted services are recovering; we’re continuing to monitor for intermittent issues
-
resolved Aug 31, 2026, 08:10 AM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 30, 2026, 10:45 PM UTC
- Resolved
- Aug 30, 2026, 10:45 PM UTC
- Duration
- —
Timeline · 1 update
-
resolved Aug 30, 2026, 10:45 PM UTC
We saw Sprite deletion jobs failing between 21:18 and 22:05 UTC. This issue has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 28, 2026, 09:56 PM UTC
- Resolved
- Aug 28, 2026, 10:36 PM UTC
- Duration
- 40m
Affected: GRU - Sao Paulo, Brazil
Timeline · 5 updates
-
investigating Aug 28, 2026, 09:56 PM UTC
We are investigating networking issues impacting some hosts in GRU (São Paulo, Brazil) region. Some apps in GRU may experience increased latency or packet loss.
-
monitoring Aug 28, 2026, 10:00 PM UTC
Networking performance in GRU has normalized and we are no longer seeing issues. We are continuing to monitor to ensure a full recovery.
-
identified Aug 28, 2026, 10:12 PM UTC
We are seeing a recurrance in networking issues in GRU. Some apps in the region may experience increased latency or packet loss. We are working with our upstream networking provider to resolve.
-
identified Aug 28, 2026, 10:36 PM UTC
Our upstream provider has implemented a fix. Network performance in GRU has normalized.
-
resolved Aug 28, 2026, 10:36 PM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 28, 2026, 08:09 AM UTC
- Resolved
- Aug 28, 2026, 10:22 AM UTC
- Duration
- 2h 13m
Affected: Customer Applications
Timeline · 2 updates
-
investigating Aug 28, 2026, 08:09 AM UTC
We are currently investigating this issue.
-
resolved Aug 28, 2026, 10:22 AM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 26, 2026, 06:14 PM UTC
- Resolved
- Aug 26, 2026, 06:43 PM UTC
- Duration
- 28m
Timeline · 4 updates
-
investigating Aug 26, 2026, 06:14 PM UTC
We are investigating issues with our WireGuard gateways. Some CLI commands like `flyctl ssh console` or `flyctl proxy` may not work at this time. Apps continue to run.
-
monitoring Aug 26, 2026, 06:27 PM UTC
A fix has been implemented and we are monitoring the results.
-
monitoring Aug 26, 2026, 06:33 PM UTC
Our testing and monitoring indicates gateways should be back to normal; if you are still having problem using `flyctl ssh console`, try restarting the `flyctl` agent by `flyctl agent restart`.
-
resolved Aug 26, 2026, 06:43 PM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 24, 2026, 10:23 AM UTC
- Resolved
- Aug 24, 2026, 01:24 PM UTC
- Duration
- 3h 1m
Affected: Metrics
Timeline · 3 updates
-
investigating Aug 24, 2026, 10:23 AM UTC
We are currently experiencing some metrics lag on servers in some regions. We are provisioning more metric processing instances to accommodate the backlog and catch up.
-
monitoring Aug 24, 2026, 12:47 PM UTC
All hosts have caught up with metrics and we're monitoring the situation
-
resolved Aug 24, 2026, 01:24 PM UTC
This is now resolved
Read the full incident report →
- Detected by Pingoru
- Aug 23, 2026, 01:28 AM UTC
- Resolved
- Aug 23, 2026, 02:10 AM UTC
- Duration
- 41m
Affected: Customer Applications
Timeline · 3 updates
-
investigating Aug 23, 2026, 01:28 AM UTC
We are investigating network issues in the Los Angeles region. Apps may experience higher latency or be unreachable at this time.
-
monitoring Aug 23, 2026, 02:04 AM UTC
Upstream networking issues have resolved.
-
resolved Aug 23, 2026, 02:10 AM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 20, 2026, 08:07 PM UTC
- Resolved
- Aug 20, 2026, 07:30 PM UTC
- Duration
- —
Timeline · 1 update
-
resolved Aug 20, 2026, 08:07 PM UTC
A BGP configuration error caused our Anycast DNS to route to some nodes without the proper DNS infrastructure. The issue was temporary and was resolved as soon as we removed that node from BGP.
Read the full incident report →
- Detected by Pingoru
- Aug 20, 2026, 01:54 PM UTC
- Resolved
- Aug 20, 2026, 02:16 PM UTC
- Duration
- 22m
Timeline · 3 updates
-
identified Aug 20, 2026, 01:54 PM UTC
We have identified an issue causing authentication errors for some operations from `flyctl`. These operations are failing with an error like: `This endpoint no longer accepts legacy OAuth tokens (starting with `fo1_`). Please use a macaroon token (starting with `fm2_`) instead. We have identified the issue and are rolling out a fix
-
monitoring Aug 20, 2026, 02:05 PM UTC
A fix has been deployed and this error should no longer be occurring. We're monitoring to ensure full recovery.
-
resolved Aug 20, 2026, 02:16 PM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 20, 2026, 07:25 AM UTC
- Resolved
- Aug 20, 2026, 07:57 AM UTC
- Duration
- 32m
Affected: Management Plane - ORD
Timeline · 4 updates
-
investigating Aug 20, 2026, 07:25 AM UTC
We had an issue with the ord-0 Fly Kubernetes cluster, and many MPG clusters are failing to restart. Our MPG team is actively working on it.
-
investigating Aug 20, 2026, 07:26 AM UTC
We are continuing to investigate this issue.
-
monitoring Aug 20, 2026, 07:32 AM UTC
A fix has been implemented and we are monitoring the results.
-
resolved Aug 20, 2026, 07:57 AM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 19, 2026, 12:00 AM UTC
- Resolved
- Aug 19, 2026, 12:00 AM UTC
- Duration
- —
Timeline · 1 update
-
resolved Aug 19, 2026, 01:15 AM UTC
6PN networking issues between some machines in YYZ during a rollout which was rolled back once we noticed errors. During this time some machines were unable to talk to internal resources like other DBs, other apps or MPG clusters.
Read the full incident report →
- Detected by Pingoru
- Aug 17, 2026, 01:19 PM UTC
- Resolved
- Aug 17, 2026, 03:41 PM UTC
- Duration
- 2h 21m
Affected: Deployments
Timeline · 2 updates
-
investigating Aug 17, 2026, 01:19 PM UTC
New machines may fail to create in ARN because we lack capacity.
-
resolved Aug 17, 2026, 03:41 PM UTC
The capacity issue in the ARN region has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 14, 2026, 07:30 PM UTC
- Resolved
- Aug 14, 2026, 08:33 PM UTC
- Duration
- 1h 3m
Affected: Machines API
Timeline · 3 updates
-
identified Aug 14, 2026, 07:30 PM UTC
We are working to recover our secrets service after a failed deployment. Apps continue to run, but it is not possible to create new apps or update secrets at this time.
-
monitoring Aug 14, 2026, 08:08 PM UTC
We have failed over the secrets database to a replica, and the Machines API appears healthy now. We are monitoring for any further issues.
-
resolved Aug 14, 2026, 08:33 PM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 13, 2026, 05:45 PM UTC
- Resolved
- Aug 13, 2026, 07:41 PM UTC
- Duration
- 1h 56m
Affected: Customer Applications
Timeline · 6 updates
-
investigating Aug 13, 2026, 05:45 PM UTC
We are currently investigating degraded ipv6 networking on a subset of hosts
-
investigating Aug 13, 2026, 05:51 PM UTC
We are continuing to investigate this issue.
-
identified Aug 13, 2026, 06:14 PM UTC
The issue has been identified and a fix is being implemented.
-
monitoring Aug 13, 2026, 06:36 PM UTC
A fix has been implemented and we are monitoring the results.
-
monitoring Aug 13, 2026, 06:37 PM UTC
We are continuing to monitor for any further issues.
-
resolved Aug 13, 2026, 07:41 PM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 09, 2026, 02:50 AM UTC
- Resolved
- Aug 09, 2026, 07:10 AM UTC
- Duration
- 4h 20m
Affected: Machines API
Timeline · 6 updates
-
investigating Aug 09, 2026, 02:50 AM UTC
We are currently investigating app-not-found errors returned by Machines API calls made shortly after creating new applications.
-
identified Aug 09, 2026, 03:50 AM UTC
We have identified the issue as failed insertions in a subset of Corrosion batches. These failures trigger retries, which can cause timeouts for other batches.
-
identified Aug 09, 2026, 04:58 AM UTC
We’ve applied a mitigation to reduce the impact from Corrosion batch insertion retries and are continuing to monitor while affected nodes catch up.
-
identified Aug 09, 2026, 06:03 AM UTC
We’ve deployed an additional mitigation to further reduce Corrosion retry pressure and are seeing improvement; we’re continuing to monitor while remaining affected nodes catch up.
-
monitoring Aug 09, 2026, 06:42 AM UTC
A fix has been implemented and we are monitoring the results
-
resolved Aug 09, 2026, 07:10 AM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 05, 2026, 08:26 PM UTC
- Resolved
- Aug 05, 2026, 08:41 PM UTC
- Duration
- 15m
Affected: Management Plane - IAD
Timeline · 2 updates
-
identified Aug 05, 2026, 08:26 PM UTC
High CPU pressure on a shared etcd instance is causing lags on MPG creation in the IAD region
-
resolved Aug 05, 2026, 08:41 PM UTC
Etcd is stable. Services are back to normal.
Read the full incident report →
- Detected by Pingoru
- Aug 04, 2026, 07:03 PM UTC
- Resolved
- Aug 05, 2026, 03:22 AM UTC
- Duration
- 8h 19m
Affected: Management Plane - GRU
Timeline · 3 updates
-
identified Aug 04, 2026, 07:03 PM UTC
New MPG clusters may fail to create in GRU because we lack capacity.
-
monitoring Aug 04, 2026, 08:59 PM UTC
We tweaked hosts to allow for more machine allocation. We'll be monitoring the region over the next hours.
-
resolved Aug 05, 2026, 03:22 AM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 04, 2026, 02:53 PM UTC
- Resolved
- Aug 04, 2026, 11:58 PM UTC
- Duration
- 9h 4m
Affected: SSL/TLS Certificate Provisioning
Timeline · 5 updates
-
investigating Aug 04, 2026, 02:53 PM UTC
We're currently investigating issues related to certificate issuance. Certificates may be delayed for new custom domains.
-
identified Aug 04, 2026, 05:02 PM UTC
We believe we have identified the issue and are releasing a fix.
-
identified Aug 04, 2026, 06:26 PM UTC
We have the size of TLS certificate issuance backlog under control, but are still seeing some remaining issues and are currently working to clean up the edge cases.
-
monitoring Aug 04, 2026, 08:36 PM UTC
A fix has been implemented and we are monitoring the results.
-
resolved Aug 04, 2026, 11:58 PM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 03, 2026, 06:50 PM UTC
- Resolved
- Aug 03, 2026, 06:50 PM UTC
- Duration
- —
Timeline · 1 update
-
resolved Aug 03, 2026, 03:16 PM UTC
A BGP misconfiguration while provisioning new edge capacity caused most traffic from Europe endpoints to be dropped, between 14:50 UTC and 15:02 UTC. The misconfiguration has been fixed and we are implementing safeguards against this kind of issue in the future.
Read the full incident report →