Cornerstone incident

Major Issue: Career Site is not unavailable for FRA SL1

Major Resolved View vendor source →

Cornerstone experienced a major incident on August 28, 2026 affecting Uptime, lasting 3h 2m. The incident has been resolved; the full update timeline is below.

Started
Aug 28, 2026, 04:11 PM UTC
Resolved
Aug 28, 2026, 07:14 PM UTC
Duration
3h 2m
Detected by Pingoru
Aug 28, 2026, 04:11 PM UTC

Affected components

Uptime

Update timeline

  1. investigating Aug 28, 2026, 04:11 PM UTC

    A production incident occurred in the EU Production environment hosted in Frankfurt (FRA PRD), resulting in the Careersite service becoming completely unavailable. This incident has been classified as a high-severity issue. Resolving this problem is our top priority, and we are working diligently to restore service as quickly as possible. Please check back periodically for additional updates, which will be posted as they become available.

  2. investigating Aug 28, 2026, 06:00 PM UTC

    We are continuing to investigate this issue.

  3. monitoring Aug 28, 2026, 06:03 PM UTC

    A fix has been implemented and we are monitoring the results.

  4. resolved Aug 28, 2026, 07:14 PM UTC

    After a period of monitoring and no recurrence, this is resolved.

  5. postmortem Sep 06, 2026, 06:24 PM UTC

    **Root Cause Analysis** ‌ **Incident Summary:** On August 28th, 2026, users in the FRA SL1 Production environment experienced intermittent issues accessing Career Site functionalities. **Impact:** Intermittent failures while accessing Career Site functionality, resulting in a degraded user experience and partial service disruption during the incident window. **Root Cause:** The issue was caused by a sudden increase in traffic that temporarily exceeded the available service capacity. During the resulting scale-up activity, newly provisioned application capacity required additional time to become fully operational and could not immediately accommodate the increased demand. The limited capacity buffer combined with the application startup time resulted in intermittent service failures until sufficient capacity was available to handle the traffic. **Resolution:** The service capacity and scaling configuration were adjusted to maintain an appropriate capacity buffer and better accommodate sudden increases in traffic. The affected service was also redeployed to restore stability. **Preventive Measures:** * Application startup and initialization processes will be optimized to reduce the time required for newly provisioned capacity to become fully operational * Baseline capacity and scaling thresholds will be reviewed and optimized to maintain sufficient buffer for sudden increases in traffic * Capacity and application health will continue to be monitored during traffic fluctuations to ensure sufficient resources are available to maintain service stability