Arista CloudVision incident

High Device Streaming Latency experienced on the Platform

Minor Resolved View vendor source →

Arista CloudVision experienced a minor incident on June 27, 2026 affecting Core Platform, lasting 10d 17h. The incident has been resolved; the full update timeline is below.

Started
Jun 27, 2026, 05:06 AM UTC
Resolved
Jul 07, 2026, 10:57 PM UTC
Duration
10d 17h
Detected by Pingoru
Jun 27, 2026, 05:06 AM UTC

Affected components

Core Platform

Update timeline

  1. investigating Jun 26, 2026, 05:27 PM UTC

    We are experiencing high device streaming latency on this affected cluster. Some impact may be user visible, such as: - A delay in processing streamed device data. Some events such as "Device stopped streaming" or "Streaming Latency Breached Threshold" may fire as a result of this. In this case, the device updates are behind schedule and will be processed with some delay during the impact window. - There may be some stale data, or sporadic timeouts viewing parts of the UI. - Some provisioning features may be delayed or timeout due to the high latency. We are currently investigating the issue and working on mitigation. Thank you for your patience.

  2. identified Jun 26, 2026, 06:10 PM UTC

    We are actively investigating as well as working on recovery.

  3. identified Jun 26, 2026, 06:34 PM UTC

    Our Publish latency for P90 and P95 have stabilized. We are continuing with recovery.

  4. monitoring Jun 26, 2026, 06:42 PM UTC

    We have fully recovered from the ingestion lag. We are continuing to monitor.

  5. identified Jun 26, 2026, 07:09 PM UTC

    Currently, lag has resolved and our publish latencies are stable - however we are seeing elevated error increase, which we are investigating.

  6. identified Jun 26, 2026, 08:27 PM UTC

    We are continuing to stablize the platform.

  7. identified Jun 26, 2026, 09:33 PM UTC

    We are continuing to stabilize the platform.

  8. identified Jun 26, 2026, 10:51 PM UTC

    We are continuing to stabilize the platform.

  9. identified Jun 27, 2026, 12:35 AM UTC

    We are continuing to stabilize the platform.

  10. identified Jun 27, 2026, 01:58 AM UTC

    We are continuing to stabilize the platform.

  11. identified Jun 27, 2026, 05:06 AM UTC

    Our team is continuing to respond to this incident. At this time, some devices may see intermittent streaming latency increases, and potentially erroneous "Device Stopped Streaming" events, but not at broad scale. We expect this to impact about < 1 % of our devices currently for cv-prod-us-central1-b. We are continuing to stabilize the platform.

  12. identified Jun 27, 2026, 06:37 AM UTC

    We are continuing to stabilize the platform. We will post updates every 3 hours to reduce spam.

  13. identified Jun 27, 2026, 09:39 AM UTC

    We are continuing to stabilize the platform. We will post updates every 3 hours to reduce spam.

  14. identified Jun 27, 2026, 12:41 PM UTC

    We are continuing to stabilize the platform. We will post updates every 3 hours to reduce spam.

  15. identified Jun 27, 2026, 03:46 PM UTC

    We are continuing to stabilize the platform. We will post updates every 3 hours to reduce spam.

  16. identified Jun 27, 2026, 06:32 PM UTC

    We are continuing to stabilize the platform. We will post updates every 3 hours to reduce spam.

  17. identified Jun 27, 2026, 09:32 PM UTC

    We are continuing to stabilize the platform. We will post updates every 3 hours to reduce spam.

  18. identified Jun 28, 2026, 12:53 AM UTC

    We are continuing to improve the stability of the platform. We have made significant progress since the start of this incident. We have seen steady stabilization since June 27th, 17:00 UTC, with platform ingest rates and streaming latency demonstrating consistent periods of stability. At this time, our team is focusing on mitigating brief periods ( < 5m ) of ingestion spikes. We are working on resolving this aspect before fully transitioning to the monitoring phase. Thank you for your continued patience.

  19. identified Jun 28, 2026, 04:03 AM UTC

    We are continuing to improve the stability of the platform. At this time we are seeing steady stabilization since June 27th, 17:00 UTC. Our team is observing the various mitigations in place currently prior to transitioning to monitoring phase.

  20. monitoring Jun 28, 2026, 05:52 AM UTC

    We have seen steady stabilization since June 27th, 17:00 UTC, and brief ingestion latencies ( < 5m ) being resolved with mitigations since June 28th, 02:00 UTC. Device streaming latency and data ingestion are holding steady at this time. We are continuing to monitor the service region cv-prod-us-central1-b.

  21. monitoring Jun 28, 2026, 11:24 AM UTC

    Device streaming latency and data ingestion are still steady. We are still monitoring the platform. The next update will be provided if the situation changes.

  22. monitoring Jun 29, 2026, 12:53 AM UTC

    We are continuing to monitor this service region. There will be no further updates to help reduce notification spam unless there is a significant change in the service region's state.

  23. identified Jun 29, 2026, 04:12 PM UTC

    We are investigating brief periods of elevated streaming latency. The platform is still operational.

  24. identified Jul 01, 2026, 06:46 PM UTC

    Our teams are still working around the clock to mitigate the brief spikes of high streaming latency and determine various contributing causes. Thank you for your patience.

  25. identified Jul 02, 2026, 03:47 AM UTC

    We continue to make progress towards identifying mitigation and causes for the brief elevated ingestion latencies seen on the platform. Thank you for your understanding and patience.

  26. identified Jul 03, 2026, 12:22 AM UTC

    We continue to make progress towards identifying mitigation and causes for the brief elevated ingestion latencies seen on the platform.

  27. identified Jul 03, 2026, 07:20 PM UTC

    We are continuing to investigate resolutions to the intermittent elevated P99 tail latency seen for data ingestion on the platform. The platform is fully operational. Thank you for your understanding and patience.

  28. identified Jul 06, 2026, 03:58 PM UTC

    We are continuing to investigate resolutions to the intermittent elevated P99 tail latency seen for data ingestion on the platform. The platform is fully operational. Thank you for your understanding and patience.

  29. resolved Jul 07, 2026, 10:57 PM UTC

    Hello CVaaS Users, This incident is closed as the original high impact on June 26th, 2026 where cv-prod-us-central1-b service region experienced a period of high ingestion latency that may have caused devices incorrectly being marked as Inactive on CloudVision, or experiencing sustained periods of elevated streaming latency also as reported by CloudVision. This original high impact has been effectively mitigated by June 27th, 2026. Since then our team has been working on fixing brief, intermittent streaming latencies (under 1 minute) that affect a small number of devices. At present, cv-prod-us-central1-b service region is fully operational. We expect users to be able to continue using Provisioning and Telemetry features as normal. No other service regions are impacted by this incident. We will follow up with a postmortem of this incident, especially regarding the impact on June 26th, 2026.

  30. postmortem Jul 20, 2026, 11:52 PM UTC

    **Incident Report: CVaaS Service Disruption - June 26th, 2026** On June 26th, 2026 10:00AM PDT - 12:30PM PDT CVaaS service region cv-prod-us-central1-b experienced significant ingestion delay into NetDL that caused the following incorrect states to be observed by users: * Devices incorrectly marked as Inactive, * Multiple inactionable “Device Stopped Streaming” events, * Multiple inactionable “High Streaming Latency” events During this time period, executing device provisioning actions may have been delayed or failed. There was no data loss. However, data ingested from actively streaming devices into CVaaS may have only become accessible after some delay. This may have impacted users accessing the UI, or other services consuming that data from CloudVision via API integrations. Between June 26, 2026 12:30PM PDT - June 27th, 12:30PM PDT, some users may have experienced the above behaviors intermittently impact their devices, but not block device provisioning activities. By June 27th, 2026, all mitigations were in place, and we transitioned to mitigating intermittent instances of elevated tail latency along the data ingestion path into NetDL. Between June 27th, 2026 and July 6th, 2026 a small subset of users of the cv-prod-us-central1-b service region may have experienced rare, but intermittent instances where CloudVision reported Event “High Streaming Latency”, potentially due to brief delays in NetDL ingestion. Users were not blocked from performing device provisioning activities. ‌ **Root Cause** Network performance degradation and instabilities in our underlying cloud provider were the ultimate root causes of delays ingesting data into the part of the system responsible for storing, processing, and updating device streaming status and streaming latency. Our mitigation activities between June 26th, 2026 and June 27th, 2026 stabilized the vast majority of the impact, but the changing underlying conditions prolonged the total duration of incident. This incident was not caused by software change\(s\), monitoring gaps, or software bugs in CloudVision as-a-Service. ‌ **Our Response** Multiple teams at Arista engaged around the clock from June 26th, 2026 until July 6th, 2026 to ensure we fully resolved this incident, as well as reviewed all details in our internal reviews. We continue to actively monitor this service region, as well as all of our CloudVision as-a-Service service regions. Arista takes full responsibility for the performance and availability of the service regardless of the underlying cloud platform conditions. We are committed to making changes in both the short and long term that will provide greater resilience to underlying provider instabilities and changing real-time conditions. Thank you for using CloudVision as-a-Service.