Zonos incident

2026-08-03 Elevated error rates for landed cost quotes

Minor Resolved View vendor source →

Zonos experienced a minor incident on August 3, 2026 affecting Landed Cost API and International Checkout and 1 more component, lasting 1h 36m. The incident has been resolved; the full update timeline is below.

Started
Aug 03, 2026, 09:00 PM UTC
Resolved
Aug 03, 2026, 10:36 PM UTC
Duration
1h 36m
Detected by Pingoru
Aug 03, 2026, 09:00 PM UTC

Affected components

Landed Cost APIInternational CheckoutShopify Duty TaxZonos Prepay (iOS)Landed Cost API (Legacy)BigCommerce Duty TaxZonos Prepay (Android)QuoterMagento Duty TaxLanded Cost API (GraphQL)

Update timeline

  1. investigating Aug 03, 2026, 09:23 PM UTC

    We are currently investigating this issue.

  2. identified Aug 03, 2026, 09:24 PM UTC

    The issue has been identified and a fix is being implemented.

  3. identified Aug 03, 2026, 09:36 PM UTC

    We are continuing to work on a fix for this issue.

  4. identified Aug 03, 2026, 09:57 PM UTC

    We are continuing to work on a fix for this issue.

  5. monitoring Aug 03, 2026, 09:57 PM UTC

    A fix has been implemented and we are monitoring the results.

  6. monitoring Aug 03, 2026, 10:36 PM UTC

    We are continuing to monitor for any further issues.

  7. resolved Aug 03, 2026, 10:36 PM UTC

    This incident has been resolved.

  8. postmortem Sep 09, 2026, 08:06 PM UTC

    ## **Incident Summary** On August 3, 2026, Zonos experienced elevated error rates affecting Landed Cost and related services. The incident began at approximately 2:52 PM MDT and was resolved at 3:51 PM MDT, resulting in approximately 59 minutes of degraded service. The incident was a partial outage: requests continued to be processed throughout the event, but a significant portion of requests relying on the affected service returned errors. ## **Impact** The primary affected service experienced an approximately 50% error rate for GraphQL requests during the incident window. As a result, customers using Landed Cost experienced intermittent request failures, while other requests continued to process successfully. Because Landed Cost relies on shared platform services, the underlying issue also resulted in elevated error rates across some related workflows. ## **Root Cause** The incident was triggered by a production deployment that began using a newly introduced data access path within our platform. Under production request volume, this access pattern generated significantly more database activity than anticipated. This exhausted available database connections within a supporting service and caused requests depending on that service to fail. The underlying issue had not manifested previously because the affected functionality had not yet been exercised at production scale. ## **Resolution** Our engineering team initially rolled back a related service deployment while investigating the source of the errors. When this did not resolve the issue, the Landed Cost deployment that activated the problematic access pattern was rolled back. Service recovered immediately following that rollback. A permanent correction was subsequently deployed to optimize the database access pattern and route the affected workload appropriately. ## **Preventive Actions** Following the incident, we reviewed both the immediate failure and how it propagated through dependent services. We are taking several actions to reduce the likelihood and potential impact of similar incidents, including: * Adding additional monitoring and alerting for database connection saturation. * Strengthening safeguards and review practices for database queries against partitioned data. * Expanding testing of newly activated data-access patterns under production-like load. * Reviewing service isolation and failure handling to reduce the impact a degraded dependency can have on other platform services. * Reviewing similar data-access patterns across the platform for the conditions identified during this incident. We apologize for the disruption this incident caused. We will continue using the findings from this event to improve the resilience and observability of the Zonos platform.