Dagster incident

Dagster UI down for multiple users

Major Resolved View vendor source →

Dagster experienced a major incident on February 27, 2025 affecting Dagster Cloud UI, lasting 37m. The incident has been resolved; the full update timeline is below.

Started
Feb 27, 2025, 10:52 AM UTC
Resolved
Feb 27, 2025, 11:29 AM UTC
Duration
37m
Detected by Pingoru
Feb 27, 2025, 10:52 AM UTC

Affected components

Dagster Cloud UI

Update timeline

  1. investigating Feb 27, 2025, 09:27 AM UTC

    We've received reports from multiple users having trouble accessing the Dagster Cloud UI. Our engineering team is currently investigating.

  2. investigating Feb 27, 2025, 10:25 AM UTC

    We are continuing to investigate this issue.

  3. investigating Feb 27, 2025, 10:51 AM UTC

    We have rolled out a fix and are monitoring service restoration. The Dagster+ UI is available again. STDERR and STDOUT tabs on runs are temporarily disabled.

  4. monitoring Feb 27, 2025, 10:52 AM UTC

    We have rolled out a fix and are monitoring service restoration. The Dagster+ UI is available again. STDERR and STDOUT tabs on runs are temporarily disabled.

  5. resolved Feb 27, 2025, 11:29 AM UTC

    The Dagster+ UI is restored. Agents, runs, and automations are not affected. We will continue to work to restore STDOUT and STDERR in the UI during US business hours. In the meantime, compute logs will continue to upload but will not be accessible via the UI. [UPDATE 1:10 pm EST] STDERR and STDOUT are available again in the UI for all users.

  6. postmortem Mar 07, 2025, 09:12 PM UTC

    In addition to [structured event logs](https://docs.dagster.io/guides/monitor/logging/#structured-event-logs) that appear in the Dagster UI, Dagster supports [raw compute logs](https://docs.dagster.io/guides/monitor/logging/#raw-compute-logs) to capture STDERR and STDOUT. During the affected window, a fault in our compute log download process resulted in timeouts and diminished responsiveness of our web servers. We mitigated the issue by temporarily disabling compute log downloads before adjusting compute log behavior so that the conditions that led to the failure are no longer possible. Log functionality has since been fully restored. We've also adjusted our incident process to ensure status pages are posted more quickly in the future. Please reach out if you have additional questions. Thank you for your patience.