Flexera incident

Spot Ocean – AWS – Ocean ECS Console Performance Degradation

Major Resolved View vendor source →

Flexera experienced a major incident on July 6, 2026 affecting Spot UI, lasting 2h 38m. The incident has been resolved; the full update timeline is below.

Started
Jul 06, 2026, 02:19 PM UTC
Resolved
Jul 06, 2026, 04:57 PM UTC
Duration
2h 38m
Detected by Pingoru
Jul 06, 2026, 02:19 PM UTC

Affected components

Spot UI

Update timeline

  1. investigating Jul 06, 2026, 02:19 PM UTC

    Incident Description: We are currently investigating degraded performance affecting the Spot Ocean ECS console for AWS customers. Customers may experience significant slowness or delays when accessing or navigating the console. Priority: P2 Restoration Activity: Our technical teams are actively investigating and working to restore normal functionality. We will provide further updates as more information becomes available.

  2. monitoring Jul 06, 2026, 03:11 PM UTC

    Our investigation identified that the degradation was related to an AWS service health event impacting the us-east-1 region. AWS has since reported that the issue has been resolved, and we are now seeing signs of recovery on our side, including improved Spot Ocean ECS console performance and successful completion of related AWS workflows. The production environment appears stable at this time. Our teams will continue monitoring for a short validation period to ensure performance remains stable before marking the incident resolved.

  3. resolved Jul 06, 2026, 04:57 PM UTC

    We have continued to monitor the environment for an extended period and services have remained stable, with no further issues observed. This incident has been resolved.

  4. postmortem Jul 20, 2026, 07:00 AM UTC

    **Description:** Spot Ocean – AWS – Ocean ECS Console Performance Degradation **Timeframe:** July 6, 2026, 05:45 AM PDT to July 6, 2026, 07:53 AM PDT ‌ **Incident Summary** ‌ On Monday, July 6, 2026, at 05:45 AM PDT , our teams identified a performance degradation affecting Spot Ocean ECS for AWS customers. During the impact window, customers experienced increased latency when accessing the Ocean ECS console, and some AWS-related operations were delayed or did not complete successfully. Technical teams immediately initiated an investigation to identify the source of the issue. As an AWS service disruption was occurring concurrently in the AWS us-east-1 region, the investigation initially considered both the external AWS event and a recently completed production deployment as potential contributing factors. Through detailed analysis, teams determined that the primary cause of the customer impact was instability introduced by the recent Gateway deployment. The resulting degradation in Gateway performance affected request processing and responsiveness within Spot Ocean ECS. While the concurrent AWS service disruption added complexity to the investigation, it was confirmed not to be the primary driver of the customer-facing impact. To restore service, teams rolled back the deployment to the previous stable version. Following the rollback, platform performance returned to expected levels and customer workflows resumed normal operation. An extended period of monitoring confirmed sustained service stability before the event was formally resolved. ‌ **Root Cause** ‌ The incident was caused by instability introduced in a Gateway deployment .The deployment resulted in degraded performance within the Gateway service, reducing its ability to efficiently process customer requests. This caused increased response times and intermittent failures for Spot Ocean ECS console operations and AWS-related workflows. Rolling back to the previous stable Gateway version restored normal platform performance and resolved the customer impact. Contributing Factors * A concurrent AWS service disruption in the us-east-1 region occurred during the same timeframe. While it was not the cause of the incident, it complicated the initial investigation and delayed identification of the underlying Gateway issue. * The Gateway rollback required additional time due to the scale of the production deployment, extending the overall recovery process. * Operations involving resource creation and updates experienced greater impact than read-only activities because of the degraded Gateway responsiveness. ‌ **Remediation Actions** ‌ * Technical teams investigated the degradation and identified a recent Gateway deployment as the primary source of the issue. * The affected Gateway deployment was rolled back to last know stable version, restoring service stability. * Platform performance and customer-facing workflows were validated following the rollback. * Technical teams completed an extended monitoring period to confirm stable service operation before closing the incident. ‌ **Future Preventative Measures** ‌ * Improved Deployment Safeguards - Engineering teams will strengthen deployment controls and validation processes to reduce the likelihood of similar deployment-related issues affecting production environments. * Enhanced Platform Monitoring - Additional monitoring and alerting will be implemented to provide earlier detection of abnormal Gateway behavior, including service responsiveness, resource utilization, and application stability. * Improved Recovery Procedures - Rollback procedures will be enhanced and regularly validated to reduce recovery time and improve operational efficiency during deployment-related incidents. * Proactive Service Health Validation - Synthetic health checks will be expanded to continuously validate critical Ocean ECS user workflows, enabling earlier identification of customer-facing performance degradation. * Enhanced Dependency Monitoring - Engineering teams will continue improving visibility into external cloud service events and platform dependencies to enable faster differentiation between internal platform issues and third-party service disruptions during future investigations.