Zaptec incident

Some users may experience delays in device status updates.

Minor Resolved View vendor source →

Zaptec experienced a minor incident on September 24, 2026 affecting OCPP and Portal, lasting 6d 19h. The incident has been resolved; the full update timeline is below.

Started
Sep 24, 2026, 11:26 AM UTC
Resolved
Oct 01, 2026, 07:15 AM UTC
Duration
6d 19h
Detected by Pingoru
Sep 24, 2026, 11:26 AM UTC

Affected components

OCPPPortal

Update timeline

  1. investigating Sep 24, 2026, 11:26 AM UTC

    We are currently working on improving cloud stability. This is an unscheduled maintenance. During this time, some users may experience delays in device status updates. Device operation is not expected to be affected, but the displayed status may take longer than usual to refresh. We apologise for the inconvenience and will provide an update when the situation has improved.

  2. investigating Sep 24, 2026, 01:12 PM UTC

    We're declaring incident. Cloud functions, but there might be some issues due to delayed processing. We will give updates around every hour.

  3. investigating Sep 24, 2026, 02:20 PM UTC

    We are still investigating the root cause, from what we see charging in Zaptec mode should not be affected, but OCPP bridge could be.

  4. investigating Sep 24, 2026, 03:33 PM UTC

    We are continuing to investigate this issue.

  5. investigating Sep 24, 2026, 04:43 PM UTC

    We are continuing to investigate this issue.

  6. monitoring Sep 24, 2026, 06:35 PM UTC

    A fix has been implemented and we are monitoring the results.

  7. monitoring Sep 25, 2026, 07:13 AM UTC

    We are continuing to monitor for any further issues.

  8. resolved Oct 01, 2026, 07:15 AM UTC

    This incident has been resolved.

  9. postmortem Oct 01, 2026, 11:07 AM UTC

    ## Incident summary On September 24, 2026, we identified a delay in one of our telemetry-processing services. The service was receiving data faster than it could process it, which created a growing backlog. Core charging functionality and primary communication services continued to operate normally. However, delayed telemetry processing affected the responsiveness of the Solar feature for some APM devices. The backlog was fully processed by October 1, 2026, and the incident was subsequently resolved. ## Customer impact The incident had a limited impact on customers: * Core charging operations remained available. * Primary device communication services were not affected. * The Solar feature for some APM devices experienced delays because it depends on timely telemetry data. * No permanent loss of service was identified. ## Incident timeline September 24, 13:13 — Monitoring detected a growing difference between the amount of telemetry received and the amount processed. The team began investigating and initiated incident communications. September 24, 15:06 — An internal incident was declared. The service appeared healthy according to the available monitoring, but its processing capacity was insufficient to keep up with incoming data. September 24, 15:19 — The team identified oversized telemetry records that could cause groups of messages to be processed repeatedly. A change was deployed to isolate and discard invalid records rather than retrying the entire group. September 24, 15:22 — The change improved resilience but did not resolve the processing delay. The investigation found that the service's limited monitoring made it difficult to identify the main performance constraint. September 24, 15:57 — Further analysis indicated that the issue was within the application's processing method rather than the underlying data-storage service. September 24, 17:12 — An updated version of the service was prepared for testing. Because core services remained operational, the team prioritized careful validation over an immediate production deployment to avoid introducing additional risk. September 24, 21:41 — The updated service was running successfully in the development environment. Testing continued to confirm that all processing scenarios worked as expected. September 30, 17:01 — The updated service was deployed to production. Processing performance improved immediately, and the telemetry backlog began decreasing. October 1, 00:30 — The telemetry backlog was fully processed. October 1, 08:30 — The incident was declared resolved. ## Root cause The affected service processed data-storage operations one at a time. As the volume of telemetry increased, this sequential approach could no longer provide sufficient throughput, causing the backlog to grow. The service was also based on an older application design with limited performance monitoring. This made the underlying bottleneck difficult to identify and extended the investigation. Oversized telemetry records contributed to unnecessary reprocessing but were not the primary cause of the incident. ## Resolution We replaced the affected service with an updated version that can safely process multiple storage operations concurrently. This substantially increased processing capacity and allowed the service to clear the backlog. We also improved the handling of oversized records so that a single invalid record no longer causes a larger group of telemetry messages to be reprocessed. ## Preventive actions To reduce the likelihood and duration of similar incidents, we have: Modernized the affected service and aligned it with our current application standards. Added improved monitoring for service health, processing performance, and backlog growth. Introduced controlled parallel processing to support higher telemetry volumes. Improved the handling of oversized or invalid telemetry records. Begun reviewing the wider telemetry-processing architecture to simplify data flows and make service dependencies and customer impact easier to identify. This was the final application using the relevant legacy architecture. Its modernization removes a significant source of technical debt and improves our ability to detect, diagnose, and resolve future performance issues.