Balena Outage History

Balena is up right now

Balena had 12 outages in the last 2 years totaling 609h 38m of downtime — averaging 0.5 incidents per month.

There were 12 Balena outages since October 20, 2025 totaling 609h 38m of downtime. Each is summarised below — incident details, duration, and resolution information.

Source: https://status.balena.io

Major September 28, 2026

Elevated Dashboard Errors

Detected by Pingoru
Sep 28, 2026, 04:23 PM UTC
Resolved
Sep 28, 2026, 06:14 PM UTC
Duration
1h 50m
Affected: APIDashboard
Timeline · 5 updates
  1. investigating Sep 28, 2026, 04:23 PM UTC

    We're experiencing an elevated level of errors on our Dashboard and are currently looking into the issue.

  2. identified Sep 28, 2026, 04:46 PM UTC

    The issue has been identified and a fix is being implemented.

  3. monitoring Sep 28, 2026, 04:50 PM UTC

    A fix has been implemented and we are monitoring the results.

  4. resolved Sep 28, 2026, 06:14 PM UTC

    This incident has been resolved.

  5. postmortem Sep 29, 2026, 02:18 PM UTC

    On September 28, the balenaCloud dashboard failed to load because some API requests stopped responding. A runtime upgrade to our API interacted badly with a legacy storage library, causing requests that depended on it to hang. We rolled back the upgrade, which restored service. To prevent this from happening again, we are removing the legacy code path involved and adding customizable timeouts for external backends.

Read the full incident report →

Minor July 27, 2026

Elevated Error on Public Device Url

Detected by Pingoru
Jul 27, 2026, 04:49 PM UTC
Resolved
Jul 27, 2026, 06:41 PM UTC
Duration
1h 52m
Affected: Device URLs
Timeline · 4 updates
  1. investigating Jul 27, 2026, 04:49 PM UTC

    We're currently seeing some Gateway error on public device url. Investigation ongoing.

  2. investigating Jul 27, 2026, 05:06 PM UTC

    We are continuing to investigate this issue.

  3. monitoring Jul 27, 2026, 05:06 PM UTC

    A fix has been implemented and we are monitoring the results.

  4. resolved Jul 27, 2026, 06:41 PM UTC

    Unusual load on the proxy, caused some degraded performances on the device URL. Everything is now working properly.

Read the full incident report →

Major July 6, 2026

Elevated API Errors

Detected by Pingoru
Jul 06, 2026, 01:47 PM UTC
Resolved
Jul 06, 2026, 02:37 PM UTC
Duration
50m
Affected: API
Timeline · 4 updates
  1. identified Jul 06, 2026, 01:47 PM UTC

    We're experiencing an elevated level of API errors and are currently looking into the issue.

  2. monitoring Jul 06, 2026, 01:53 PM UTC

    A fix has been implemented and we are monitoring the results.

  3. resolved Jul 06, 2026, 02:37 PM UTC

    This incident has been resolved.

  4. postmortem Jul 09, 2026, 09:08 PM UTC

    A routine version upgrade to a backend cache changed how the API handled its request queues. Under unusually heavy request load, those queues filled up and the API started dropping traffic, returning timeouts and errors. Over the course of the incident, affected users saw elevated API error rates that intermittently disrupted dashboard logins, API access, and device connectivity. We mitigated it by rate-limiting the abnormal traffic at our edge and rolling the cache back to the previous version, which restored normal queue behaviour and recovered the platform. To avoid a repeat, we're improving rate limiting so a single heavy consumer can't degrade service for others, and exploring autoscaling and load testing so infrastructure changes are validated under realistic load before they reach production.

Read the full incident report →

Minor March 31, 2026

Elevated GIT/Application Builder Errors

Detected by Pingoru
Mar 31, 2026, 12:55 PM UTC
Resolved
Apr 21, 2026, 04:30 PM UTC
Duration
21d 3h
Affected: Application Builder
Timeline · 5 updates
  1. identified Mar 31, 2026, 12:55 PM UTC

    We're experiencing an elevated level of errors in our application builder infrastructure and are currently looking into the issue.

  2. monitoring Apr 08, 2026, 07:51 PM UTC

    A fix has been implemented and we are monitoring the results.

  3. monitoring Apr 08, 2026, 07:51 PM UTC

    We are continuing to monitor for any further issues.

  4. resolved Apr 21, 2026, 04:30 PM UTC

    This incident has been resolved.

  5. postmortem Apr 21, 2026, 05:06 PM UTC

    Starting around March 11, some cloud builds began failing intermittently with no such image errors. The failures were non-deterministic and affected all architectures. At peak, some users saw around 50% failure rates. We identified and fixed several bugs in the builder's image garbage collector that caused it to over-count freed disk space and run too aggressively, eventually deleting images that in-progress builds still needed. Fixes were deployed between March 19 and April 14, with build failure rates dropping to near-zero after the final deploy. We're continuing to monitor and working on additional safeguards to prevent the garbage collector from targeting images that active builds depend on.

Read the full incident report →

Minor March 23, 2026

Builder Degraded performance

Detected by Pingoru
Mar 23, 2026, 04:57 PM UTC
Resolved
Mar 25, 2026, 01:18 PM UTC
Duration
1d 20h
Affected: Application Builder
Timeline · 4 updates
  1. investigating Mar 23, 2026, 04:57 PM UTC

    We are seeing several builds intermitently failing with 404 errors - No such image during builds and are investigating.

  2. monitoring Mar 23, 2026, 06:55 PM UTC

    A fix has been implemented and we are monitoring the results.

  3. resolved Mar 25, 2026, 01:18 PM UTC

    This incident has been resolved.

  4. postmortem Mar 25, 2026, 03:38 PM UTC

    Between March 11 and March 25, some cloud builds experienced intermittent failures with "no such image" errors. The issue was non-deterministic and did not affect all builds. We've identified a likely contributing factor and deployed mitigations that have stabilized build reliability. We're continuing to investigate the underlying cause to prevent recurrence. If you experienced build failures during this window, re-running your build should succeed. We appreciate your patience while we worked through this, and we apologize for the disruption.

Read the full incident report →

Major February 10, 2026

Elevated Delta Errors

Detected by Pingoru
Feb 10, 2026, 09:33 AM UTC
Resolved
Feb 11, 2026, 12:02 AM UTC
Duration
14h 28m
Affected: Delta Image Downloads
Timeline · 4 updates
  1. investigating Feb 10, 2026, 09:33 AM UTC

    Some delta generation requests are encountering errors and failing. We are currently investigating this issue.

  2. monitoring Feb 10, 2026, 10:25 AM UTC

    We have identified the potential cause and have rolled back the changes.

  3. resolved Feb 11, 2026, 12:02 AM UTC

    This incident has been resolved.

  4. postmortem Feb 11, 2026, 12:10 AM UTC

    v2 delta generation service experienced failures from ~21:15 UTC Feb 9 to ~10:00 UTC Feb 10, 2026, due to a missing configuration dependency during a logic change. **Impact:** * v2 delta generation requests failed to complete * No data loss or security impact **Root Cause:** Recent logic changes were deployed without the required accompanying configuration update, preventing the service from completing v2 delta requests. **Resolution:** The logic changes were rolled back, restoring the service to its previous stable state. **Follow-up Actions:** * Prepare and deploy the permanent fix We apologize for the disruption and any inconvenience this caused. We are committed to improving our processes to prevent similar issues in the future.

Read the full incident report →

Critical February 10, 2026

Elevated Cloudlink Errors

Detected by Pingoru
Feb 10, 2026, 04:05 AM UTC
Resolved
Feb 10, 2026, 07:56 AM UTC
Duration
3h 51m
Affected: Cloudlink (VPN)
Timeline · 5 updates
  1. investigating Feb 10, 2026, 04:05 AM UTC

    We're experiencing an elevated level of errors in our Cloudlink infrastructure and are currently looking into the issue.

  2. identified Feb 10, 2026, 06:39 AM UTC

    The issue has been identified and a fix is being implemented.

  3. monitoring Feb 10, 2026, 07:17 AM UTC

    A fix has been implemented and we are monitoring the results.

  4. resolved Feb 10, 2026, 07:56 AM UTC

    This incident has been resolved.

  5. postmortem Feb 10, 2026, 11:52 AM UTC

    Balena devices were unable to connect to Cloudlink on February 10, 2026, from approximately 02:26 GMT to 07:11 GMT due to an expired server certificate. Devices that were already connected to Cloudlink were unaffected unless the connection was terminated. **Root Cause:** The Cloudlink servers were using an expired certificate that was due for replacement. Consequently, incoming Cloudlink connections failed with a certificate verification error. **Resolution:** The certificate has been replaced, and Cloudlink servers were restarted to use the new certificate. Balena devices are expected to reconnect to Cloudlink within a few minutes after being disconnected due to the restart. **Follow-up Actions:** * Expand certificate expiry monitoring coverage to include all active certificates * Automate the certificate renewal process for Cloudlink We apologize for any disruption this caused and appreciate your patience as we continue improving our processes and operations.

Read the full incident report →

Minor December 23, 2025

Degraded Performance

Detected by Pingoru
Dec 23, 2025, 04:32 PM UTC
Resolved
Dec 24, 2025, 09:38 AM UTC
Duration
17h 6m
Affected: APIApplication Builderbalenahub
Timeline · 3 updates
  1. investigating Dec 23, 2025, 04:32 PM UTC

    We are currently investigating an issue affecting the availability of balenaCloud services.

  2. monitoring Dec 23, 2025, 07:03 PM UTC

    Scaling issues during service deployment caused by unavailability of nodes from underlying scaling provider.

  3. resolved Dec 24, 2025, 09:38 AM UTC

    Insufficient AWS compute capacity overloaded the remaining nodes. This high load caused readiness probes to fail, triggering API restarts that created a feedback loop of increasing pressure.

Read the full incident report →

Critical December 5, 2025

An upstream provider outage is affecting connectivity to balenaCloud services

Detected by Pingoru
Dec 05, 2025, 09:05 AM UTC
Resolved
Dec 05, 2025, 09:37 AM UTC
Duration
31m
Affected: APIDashboardWebsite
Timeline · 3 updates
  1. identified Dec 05, 2025, 09:05 AM UTC

    CloudFlare, our proxy provider, is having service issues. Connectivity to balenaCloud services are currently affected.

  2. monitoring Dec 05, 2025, 09:16 AM UTC

    Our upstream provider has implemented some fixes. balenaCloud services are back online. We are still monitoring the situation.

  3. resolved Dec 05, 2025, 09:37 AM UTC

    This incident has been resolved.

Read the full incident report →