Ashby incident

Ashby Unavailable

Major Resolved View vendor source →

Ashby experienced a major incident on June 1, 2026 affecting Job Feed and Google and 1 more component, lasting 2h 30m. The incident has been resolved; the full update timeline is below.

Started
Jun 01, 2026, 06:33 AM UTC
Resolved
Jun 01, 2026, 09:04 AM UTC
Duration
2h 30m
Detected by Pingoru
Jun 01, 2026, 06:33 AM UTC

Affected components

Job FeedGoogleAshby APIRecruitingSendGrid API v3Office 365Reports APIAnalyticsGoogle GmailSendGrid Parse API

Update timeline

  1. investigating Jun 01, 2026, 06:33 AM UTC

    We are currently investigating this issue.

  2. investigating Jun 01, 2026, 06:33 AM UTC

    We are continuing to investigate this issue.

  3. investigating Jun 01, 2026, 06:34 AM UTC

    We are currently investigating the issue

  4. investigating Jun 01, 2026, 06:49 AM UTC

    We are continuing to investigate this issue.

  5. identified Jun 01, 2026, 07:15 AM UTC

    We are seeing signs of recovery, though there may be residual instability. AI features will be unavailable as we mitigate the remaining issues.

  6. monitoring Jun 01, 2026, 07:31 AM UTC

    The application has recovered. AI features are still affected, and we are actively working to mitigate the issues.

  7. monitoring Jun 01, 2026, 08:07 AM UTC

    AI features are enabled again, and we're processing the backlog of accumulated work.

  8. monitoring Jun 01, 2026, 08:29 AM UTC

    All systems are now operational, we are continuing to monitor

  9. resolved Jun 01, 2026, 09:04 AM UTC

    This incident has been resolved.

  10. postmortem Jun 19, 2026, 06:19 PM UTC

    ## **Summary** On Sunday, May 31st, at 10pm PST, Ashby was slow and had a high error rate for approximately 1 hour, from 10pm to 11:10pm \(PST\). During this time, the Ashby main application would have been slow, and approximately 50% of user actions would have resulted in Ashby displaying an error. Some candidates may have noticed a slower response to their job applications, but all applications were successful and none were lost. ## **Why did this happen?** On Sunday, May 31st, at 4pm PST, an account with one of our AI providers became unavailable. As a result, calls to that provider started to error. These calls are configured to automatically retry periodically until they succeed, but the retries caused a steady increase in the number of such calls per minute. In turn, this drained the capacity of a shared subsystem that manages many different processes at Ashby. In particular, that shared subsystem is used for storing user session data. As the shared subsystem became overwhelmed, user sessions became slower to respond. This caused requests to our web servers to take significantly longer to process. Which then caused our web servers to begin running out of capacity. At 6am, capacity reached a critical low. At this point, our system began rejecting user actions. This is done to prevent complete system failure. ## **How did we resolve the situation?** At 10:04pm PST, our on-call engineers were notified. By 10:16pm PST, our on-call engineers were beginning an investigation. At 11:03pm PST, we identified the subsystem that had reached capacity. At 11:07pm PST, we stopped processing certain types of automated actions. At 11:09pm PST, we identified the issue with the external provider and rectified it immediately. At 11:10pm PST, error rates were down to 10%. At 11:20pm PST, error rates were down to normal levels \(~0%\). At 11:54pm PST, normal service was resumed, and no data was lost. ## **What have we put in place to prevent it from happening in the future?** We’ve done an internal postmortem on this incident and have implemented or plan to implement a variety of changes that achieve three things: 1. Reduced the likelihood of future failure 2. Faster incident response times 3. Lower impact on customers in the event of failure Specifically, we have done the following: * We were already in the process of replacing the subsystem at the center of the outage. We have accelerated that work. The replacement is not vulnerable to this particular kind of failure. * We’ve added additional monitoring around our AI providers to alert us to failures like this sooner. * We are prioritizing a fix for the retry mechanism that led to this incident. * We are migrating our user session storage to a new subsystem that is not vulnerable to this particular kind of failure.