Flexera incident

Snow Atlas - APAC - Login Failures and 404 Errors

Critical Resolved View vendor source →

Flexera experienced a critical incident on July 27, 2026 affecting Snow Atlas - Australia and Snow Atlas API - Australia, lasting 1h 49m. The incident has been resolved; the full update timeline is below.

Started
Jul 27, 2026, 12:11 AM UTC
Resolved
Jul 27, 2026, 02:00 AM UTC
Duration
1h 49m
Detected by Pingoru
Jul 27, 2026, 12:11 AM UTC

Affected components

Snow Atlas - AustraliaSnow Atlas API - Australia

Update timeline

  1. investigating Jul 27, 2026, 12:11 AM UTC

    Incident Description: We are investigating an issue affecting Snow Atlas customers in the APAC region. Affected users may be unable to access the platform and may encounter 404 errors when attempting to log in. Current investigation indicates the issue is limited to the APAC region, with no impact confirmed in other regions at this time. Priority: P1 Restoration Activity: Our technical teams are actively investigating the issue to identify the underlying cause and restore service. We are monitoring the environment closely and implementing corrective actions as appropriate. Further updates will be provided as more information becomes available.

  2. investigating Jul 27, 2026, 02:00 AM UTC

    Our teams have resolved the issue affecting Snow Atlas customers in the APAC region. The issue occurred when a platform service failed to start correctly, resulting in intermittent errors and access issues for some customers. Technical teams identified the underlying startup failure and restored service by restarting the affected platform components. Validation activities have confirmed successful access to the platform, and we are not observing any ongoing customer impact at this time. Our teams will continue to monitor the environment closely to ensure continued stability.

  3. resolved Jul 27, 2026, 02:00 AM UTC

    This incident has been resolved.

  4. postmortem Aug 07, 2026, 08:40 AM UTC

    **Description:** Snow Atlas - APAC - Login Failures and 404 Errors **Timeframe:** July 26, 2026, 5:00 PM PDT to July 26, 2026, 6:15 PM PDT ‌ **Incident Summary** ‌ On Sunday, 26 July 2026, Flexera experienced a service disruption affecting Snow Atlas customers in the APAC Production environment, resulting in login failures and temporary inability to access Agreement pages. The issue was isolated to the APAc region, with no impact to other production regions. Investigation by our technical teams determined that the disruption was caused by a backend service communication issue that prevented Agreement data from being retrieved successfully. The condition was consistent with a runtime synchronization issue following recent platform infrastructure maintenance. Flexera teams validated platform health, restored the affected services, and confirmed successful recovery. Customer access was fully restored, and post-recovery monitoring confirmed stable operations with no recurrence of the underlying service communication errors. A detailed post-incident review was completed, and corrective measures have been implemented to further strengthen platform resilience and recovery processes. ‌ **Root Cause** ‌ The incident was caused by an application communication failure within the APAC Production environment. A component responsible for processing requests between internal application services did not successfully handle Agreement-related requests, preventing those requests from being completed and resulting in customer-facing 404 errors and failures accessing Agreement pages. ‌ **Contributing Factors** ‌ * The application communication failure occurred following a recent infrastructure upgrade to the APAC Production environment. * Although the application services and underlying infrastructure remained healthy, an internal service responsible for processing Agreement requests did not fully recover after the upgrade. * Existing health checks validated application and infrastructure availability but did not verify the successful initialization of the internal communication path following the upgrade. * The issue was therefore not detected until customers began experiencing login failures and HTTP 404 errors when accessing Agreement pages. ‌ **Remediation Actions** ‌ To restore service, technical teams: * Confirmed the underlying infrastructure remained healthy throughout the incident. * Identified the failed internal service communication affecting Agreement requests. * Restarted the affected application services. * Validated successful recovery across impacted tenants. * Continued monitoring to confirm service stability. ‌ **Future Preventative Measures** ‌ * Enhanced Post-Upgrade Validation: Introduce additional validation checks following infrastructure upgrades to verify that critical application components and internal service communications are functioning as expected. * Improved Monitoring and Alerting: Enhance monitoring and alerting to detect internal service communication failures and responder registration issues more quickly. * Synthetic Health Checks: Implement synthetic health checks that validate critical customer workflows, including Agreement page functionality, following infrastructure changes.