Cornerstone incident

Major Issue: Career Site is not unavailable for FRA SL1

Major Resolved View vendor source →

Cornerstone experienced a major incident on September 3, 2026 affecting Response Time, lasting 4h 19m. The incident has been resolved; the full update timeline is below.

Started
Sep 03, 2026, 11:52 AM UTC
Resolved
Sep 03, 2026, 04:12 PM UTC
Duration
4h 19m
Detected by Pingoru
Sep 03, 2026, 11:52 AM UTC

Affected components

Response Time

Update timeline

  1. identified Sep 03, 2026, 11:52 AM UTC

    A production incident occurred in the EU Production environment hosted in Frankfurt (FRA PRD), resulting in the Careersite service becoming completely unavailable. This incident has been classified as a high-severity issue. Resolving this problem is our top priority, and we are working diligently to restore service as quickly as possible.

  2. monitoring Sep 03, 2026, 12:03 PM UTC

    Career Site is now accessible. We'll keep monitoring the portals.

  3. monitoring Sep 03, 2026, 12:04 PM UTC

    We are continuing to monitor for any further issues.

  4. resolved Sep 03, 2026, 04:12 PM UTC

    This incident has been resolved.

  5. postmortem Sep 21, 2026, 05:22 AM UTC

    **Root Cause Analysis** ‌ **Incident Summary:** Between Sep. 3rd and Sep. 8th, 2026, there were three Production incidents that affected Career Site functionality. This issue was intermittent across the FRA SL1 Production environment. These incidents occurred * Sep. 3rd from 4:37 am - 5:02 am PT * Sep. 7th from 2:13 pm - 2:25 pm PT * Sep. 8th from 6:20 am - 8:14 am PT **Impact:** Users experienced intermittent failures while accessing Career Site functionality. This resulted in a degraded user experience and partial service disruption during the incident window. **Root Cause:** The issue was caused by a sudden increase in traffic that temporarily exceeded the available service capacity. During the resulting scale-up activity, newly provisioned application capacity required additional time to become fully operational and could not immediately accommodate the increased demand. The limited capacity buffer combined with the application startup time resulted in intermittent service failures until sufficient capacity was available to handle the traffic. **Resolution:** * The service capacity and scaling configuration were adjusted to maintain an appropriate capacity buffer and better accommodate sudden increases in traffic. The affected service was also redeployed to restore stability * As a permanent improvement, the service has been moved to infrastructure with enhanced scaling capabilities to respond more rapidly to sudden increases in traffic. Since Sep. 12th, the new infrastructure has been fully handling service traffic **Preventive Measures:** * The service has been transitioned to infrastructure with improved scaling capabilities to better accommodate unexpected traffic increases * Capacity and application health will continue to be monitored during traffic fluctuations to ensure sufficient resources are available to maintain service stability