Northflank Outage History

Northflank is up right now

Northflank had 15 outages in the last 2 years totaling 17h 58m of downtime — averaging 0.6 incidents per month.

There were 15 Northflank outages since December 11, 2025 totaling 17h 58m of downtime. Each is summarised below — incident details, duration, and resolution information.

Source: https://status.northflank.com

Minor September 25, 2026

Degraded SSO login via WorkOS

Detected by Pingoru
Sep 25, 2026, 03:15 PM UTC
Resolved
Sep 25, 2026, 03:15 PM UTC
Duration
—
Timeline · 1 update
  1. resolved Sep 25, 2026, 03:15 PM UTC

    Type: Incident Duration: 46 minutes Affected Components: Northflank App Sep 25, 15:15:00 GMT+0 - Investigating - We are currently investigating this incident. Sep 25, 16:01:05 GMT+0 - Resolved - This incident has been resolved.

Read the full incident report →

Minor September 4, 2026

Addon expose/unexpose temporarily disabled in US - Central

Detected by Pingoru
Sep 04, 2026, 06:27 PM UTC
Resolved
Sep 04, 2026, 06:27 PM UTC
Duration
—
Timeline · 1 update
  1. resolved Sep 04, 2026, 06:27 PM UTC

    Type: Incident Duration: 2 hours and 25 minutes Affected Components: Addons Sep 4, 18:27:54 GMT+0 - Monitoring - We have temporarily disabled addon expose/unexpose in our US - Central region. This is due to an ongoing incident with our cloud service provider. We will re-enable this feature once we confirm that it is safe to do so. Sep 4, 20:52:58 GMT+0 - Resolved - This incident has been resolved.

Read the full incident report →

Minor August 5, 2026

Workloads Failing to Start Across Multiple Regions

Detected by Pingoru
Aug 05, 2026, 01:27 PM UTC
Resolved
Aug 05, 2026, 02:27 PM UTC
Duration
1h
Affected: Northflank Platform (Services)Northflank Platform (Addons)Northflank Platform (Jobs)
Timeline · 3 updates
  1. investigating Aug 05, 2026, 01:27 PM UTC

    We are seeing partial start up issues for workloads across multiple regions. We are currently investigating this incident.

  2. monitoring Aug 05, 2026, 01:42 PM UTC

    We have identified the issue and applied a mitigation. Previously stuck workloads are now starting successfully across affected regions. We are monitoring recovery and will confirm once all workloads are fully operational.

  3. resolved Aug 05, 2026, 02:27 PM UTC

    All affected workloads are now fully operational.

Read the full incident report →

Minor July 22, 2026

Investigating node stability issues in the US - East region

Detected by Pingoru
Jul 22, 2026, 12:17 PM UTC
Resolved
Jul 22, 2026, 01:39 PM UTC
Duration
1h 22m
Affected: Northflank Platform (Services)Northflank Platform (Addons)Northflank Platform (Jobs)
Timeline · 3 updates
  1. investigating Jul 22, 2026, 12:17 PM UTC

    We have noticed increased failure rates of nodes in the US - East region and are investigating the issue.

  2. monitoring Jul 22, 2026, 01:11 PM UTC

    We are seeing recovery across compute and storage. We are working to get the remaining workloads back online.

  3. resolved Jul 22, 2026, 01:39 PM UTC

    We have recovered affected workloads in the region and there are no more on-going node failures. The root cause was an outage on the Google Cloud Platform in us-east4 which hosts Northflank's US - East region.

Read the full incident report →

Minor July 21, 2026

Active DDoS Attack in US - Central

Detected by Pingoru
Jul 21, 2026, 11:09 AM UTC
Resolved
Jul 21, 2026, 01:58 PM UTC
Duration
2h 49m
Affected: Northflank Platform (Networking)
Timeline · 6 updates
  1. identified Jul 21, 2026, 11:09 AM UTC

    We are currently experiencing a DDoS attack in the US - Central region. This is affecting public ingress into the cluster. Internal traffic is not affected. The team is working on additional mitigations.

  2. monitoring Jul 21, 2026, 11:20 AM UTC

    We have implemented mitigations are monitoring the situation.

  3. resolved Jul 21, 2026, 11:41 AM UTC

    The DDoS attack has stopped, we will continue to monitor the affected region.

  4. identified Jul 21, 2026, 01:23 PM UTC

    We are currently experiencing a DDoS attack in the US - Central region. This is affecting public ingress into the cluster. Internal traffic is not affected. The team is working on additional mitigations.

  5. monitoring Jul 21, 2026, 01:43 PM UTC

    We have implemented mitigations are monitoring the situation.

  6. resolved Jul 21, 2026, 01:58 PM UTC

    This incident has been resolved. We will continue to monitor the affected region.

Read the full incident report →

Minor June 18, 2026

Let's Encrypt Certificate Generation Issues

Detected by Pingoru
Jun 18, 2026, 06:55 PM UTC
Resolved
Jun 18, 2026, 11:36 PM UTC
Duration
4h 40m
Affected: Northflank Platform (Certificates)
Timeline · 3 updates
  1. investigating Jun 18, 2026, 06:55 PM UTC

    We are currently observing delays in certificate generation due to an ongoing incident with Let's Encrypt: This will affect provisioning of new addons with TLS enabled, the provisioning of BYOC clusters as well as domain certification provisioning.

  2. monitoring Jun 18, 2026, 11:02 PM UTC

    No recent API errors experienced and addon provisioning is working as expected. We're still closely monitoring Let's Encrypt API responses.

  3. resolved Jun 18, 2026, 11:36 PM UTC

    No more certificate creation errors have been observed from Let's Encrypt. This incident has been resolved.

Read the full incident report →

Major June 10, 2026

Issues accessing app.northflank.com

Detected by Pingoru
Jun 10, 2026, 03:58 PM UTC
Resolved
Jun 10, 2026, 04:14 PM UTC
Duration
16m
Affected: Northflank App
Timeline · 3 updates
  1. investigating Jun 10, 2026, 03:58 PM UTC

    We are currently investigating this incident.

  2. monitoring Jun 10, 2026, 04:03 PM UTC

    We implemented a fix and are currently monitoring the result. Users should now be able to log in and access the UI. Running workloads were unaffected.

  3. resolved Jun 10, 2026, 04:14 PM UTC

    This incident has been resolved.

Read the full incident report →

Minor June 10, 2026

Europe - West (London): Intermittent connectivity issues impacting Postgres cluster management

Detected by Pingoru
Jun 10, 2026, 08:00 AM UTC
Resolved
Jun 10, 2026, 02:23 PM UTC
Duration
6h 22m
Affected: Northflank Platform (Addons)
Timeline · 5 updates
  1. investigating Jun 10, 2026, 08:00 AM UTC

    We are currently investigating this incident. You may see logs in Postgres with the following error message: ``` ERROR: Error communicating with DCS ```

  2. monitoring Jun 10, 2026, 09:59 AM UTC

    We have observed a recovery starting about 20 minutes ago. Currently, Postgres addons are stable.

  3. identified Jun 10, 2026, 10:19 AM UTC

    We saw a period of recovery between 09:45 and 10:04 UTC, but the issue has since recurred. This is tied to an ongoing incident affecting our cloud service provider in the region. We are continuing to monitor and will share further updates as we have them

  4. monitoring Jun 10, 2026, 11:53 AM UTC

    We have not observed any further issues since 11:05 UTC. We are awaiting confirmation from our cloud provider that the issue is fully resolved on their side, and we will update this status accordingly.

  5. resolved Jun 10, 2026, 02:23 PM UTC

    This incident has been resolved.

Read the full incident report →

Minor May 13, 2026

Europe - West (London): New resources not able to be created

Detected by Pingoru
May 13, 2026, 03:00 PM UTC
Resolved
May 13, 2026, 04:27 PM UTC
Duration
1h 26m
Affected: Northflank Platform (Services)Northflank Platform (Addons)Northflank Platform (Jobs)
Timeline · 3 updates
  1. investigating May 13, 2026, 03:00 PM UTC

    We are currently investigating this incident. Existing running workloads are unaffected.

  2. monitoring May 13, 2026, 03:33 PM UTC

    We are seeing an improvement; workloads are starting to be created. Connectivity has been restored, and we are continuing to monitor the situation.

  3. resolved May 13, 2026, 04:27 PM UTC

    This incident has been resolved.

Read the full incident report →

Minor May 10, 2026

Active DDoS Attack in Europe - West - Netherlands

Detected by Pingoru
May 10, 2026, 06:30 PM UTC
Resolved
May 10, 2026, 06:30 PM UTC
Duration
—
Timeline · 1 update
  1. resolved May 10, 2026, 06:30 PM UTC

    Type: Incident Duration: 2 hours and 2 minutes Affected Components: Addons, Services May 10, 18:30:00 GMT+0 - Identified - We are currently experiencing a large DDoS attack in the Europe - West - Netherlands region. This is affecting public ingress into the cluster. Internal traffic is not affected. The team is working on additional mitigations. May 10, 20:10:07 GMT+0 - Monitoring - We have put in place mitigations and are monitoring the affected region. May 10, 20:31:56 GMT+0 - Resolved - This incident has been resolved.

Read the full incident report →

Minor April 8, 2026

Addon and volume availability issues in US - Central

Detected by Pingoru
Apr 08, 2026, 06:15 AM UTC
Resolved
Apr 08, 2026, 06:15 AM UTC
Duration
—
Timeline · 1 update
  1. resolved Apr 08, 2026, 06:15 AM UTC

    Type: Incident Duration: 52 minutes Affected Components: Addons, Services Apr 8, 06:15:00 GMT+0 - Identified - We are aware of an issue with a node affecting some stateful workloads in US - Central. We are working on a fix for this incident. Apr 8, 07:07:27 GMT+0 - Resolved - We have recovered the node, and all workloads are now running again.

Read the full incident report →

Minor March 16, 2026

App UI is experiencing loading issues

Detected by Pingoru
Mar 16, 2026, 05:42 PM UTC
Resolved
Mar 16, 2026, 05:42 PM UTC
Duration
—
Timeline · 1 update
  1. resolved Mar 16, 2026, 05:42 PM UTC

    Type: Incident Duration: 34 minutes Affected Components: Northflank App Mar 16, 18:16:28 GMT+0 - Resolved - This incident has been resolved. Mar 16, 18:15:58 GMT+0 - Monitoring - We identified the root cause as a system component causing disproportionate database load. This is affecting read performance leading to degraded performance for the app UI. We have implemented a fix and are monitoring the situation. Mar 16, 17:42:54 GMT+0 - Identified - We are investing an incident where the UI is failing to load correctly.

Read the full incident report →

Minor March 9, 2026

Degraded performance in US Central region

Detected by Pingoru
Mar 09, 2026, 12:06 PM UTC
Resolved
Mar 09, 2026, 12:06 PM UTC
Duration
—
Timeline · 1 update
  1. resolved Mar 09, 2026, 12:06 PM UTC

    Type: Incident Duration: 29 minutes Affected Components: Networking, Addons, Services Mar 9, 12:06:32 GMT+0 - Investigating - We are currently investigating this incident. Mar 9, 12:35:56 GMT+0 - Resolved - The root cause has been identified as an issue with the GCP infrastructure control plane and has been resolved. There was minimal impact to running user workloads. The main impact was a delay in provisioning net new or redeploying existing workloads.

Read the full incident report →

Minor February 19, 2026

Issues with builds using the Heroku 24 Buildpack

Detected by Pingoru
Feb 19, 2026, 03:00 PM UTC
Resolved
Feb 19, 2026, 03:00 PM UTC
Duration
—
Timeline · 1 update
  1. resolved Feb 19, 2026, 03:00 PM UTC

    Type: Incident Duration: 1 hour and 34 minutes Affected Components: Builds Feb 19, 15:00:00 GMT+0 - Identified - We are aware of an issue where builds fail to start when using the Heroku 24 buildpack image. We are currently working on identifying a solution. We will provide an update when the fix has been released. Feb 19, 16:34:09 GMT+0 - Resolved - We have released a fix for the issue.

Read the full incident report →

Minor December 11, 2025

Node infrastructure stability - London Region

Detected by Pingoru
Dec 11, 2025, 01:54 PM UTC
Resolved
Dec 11, 2025, 01:54 PM UTC
Duration
—
Timeline · 1 update
  1. resolved Dec 11, 2025, 01:54 PM UTC

    Type: Incident Duration: 39 minutes Affected Components: Networking, , Addons, Jobs, Services, Northflank Platform → Dec 11, 13:54:32 GMT+0 - Identified - There are currently issues with node stability in the London region leading to partial outages for some workloads. The team has identified the issue and is implementing a mitigation. Dec 11, 14:13:28 GMT+0 - Monitoring - The mitigation is in place and the region is operating as expected. Dec 11, 14:33:03 GMT+0 - Resolved - The incident has been resolved and the cause identified. A node image security release lead to the host filesystem going into a read only mode when interacting with a specific set of workloads causing the node to become unresponsive. We have implemented a mitigation and are working on a permanent solution.

Read the full incident report →