ShipHawk incident

WMS slowness

Notice Resolved View vendor source →

ShipHawk experienced a notice incident on August 25, 2026 affecting WMS, lasting 3h 30m. The incident has been resolved; the full update timeline is below.

Started
Aug 25, 2026, 02:10 PM UTC
Resolved
Aug 25, 2026, 05:40 PM UTC
Duration
3h 30m
Detected by Pingoru
Aug 25, 2026, 02:10 PM UTC

Affected components

WMS

Update timeline

  1. investigating Aug 25, 2026, 02:10 PM UTC

    Some customers are experiencing slowness with the WMS access. Our team is actively investigating the issue.

  2. monitoring Aug 25, 2026, 04:01 PM UTC

    A fix has been implemented and we are monitoring the results.

  3. resolved Aug 25, 2026, 05:40 PM UTC

    This incident has been resolved. We will share additional details within 5 business days.

  4. postmortem Aug 27, 2026, 10:21 PM UTC

    ## Post-Incident Report: WMS Slowness and Errors for Warehouses on One Database Instance - August 25, 2026 **Status:** Resolved **Incident window:** August 25, 2026, ~04:30 - 08:45 PDT \(07:30 - 11:45 ET\) **Affected:** Customers whose databases are hosted on one WMS database instance in our US-East warehouse group: page load times of up to 20x normal, and intermittent errors on scanner and admin screens during the worst period. **Not affected:** All other WMS database instances and warehouse groups, the TMS platform, data integrity \(no transaction was lost, duplicated, or partially applied\), and all completed work - every transaction that was accepted was processed correctly. ### Were you affected? Impact was limited to customers whose WMS databases are hosted on one specific database instance in our US-East warehouse group, and only during the morning of August 25 \(approximately 04:30 - 08:45 PDT / 07:30 - 11:45 ET\). If you did not see slow page loads in the WMS during that window, your environment was not involved. All other WMS database instances and warehouse groups, and the entire TMS platform, operated normally throughout. ### Summary On the evening of August 24, a third-party ETL service that copies ShipHawk data into our data warehouse restarted a large batch of its sync jobs at once against one of our production databases. Each restarted job began re-reading a backlog of historical change data at full speed, in parallel and without any rate limiting. Read volume climbed to roughly five times the expected peak and consumed most of the disk bandwidth available to that database instance. This began overnight, when warehouse activity is light, so it had no effect on operations at the time. When US-East morning shifts started on August 25, normal WMS activity was added on top of the already-saturated channel and the instance reached its bandwidth limit. From the application's point of view this appeared as slow database responses: queries that normally return in milliseconds took far longer, work queued up behind them, and pages loaded slowly. Where a database response exceeded the application's timeout, the page returned an error instead of loading. The service remained operational throughout. Across the affected instance, customers processed approximately 70% of the volume normally handled in that window \(picks, packs and inventory moves\), though the experience was slow and at times difficult to work with, and the degree of impact varied between customers. The situation resolved by revoking the ETL service database credentials entirely at 08:37 PDT. The database recovered within eight minutes. Delayed work flushed through over the following two hours, and all affected warehouses were back to normal pace. No customer action was or is required, and no data was affected. No transaction was lost, duplicated, or partially applied - we verified this against integration logs covering the full incident window. Work that was submitted either completed correctly or failed cleanly before making any change. This is covered in more detail below. ### What was affected Impact was limited to customers whose databases are hosted on the affected WMS database instance in the US-East warehouse group. Throughout the incident window the system was slow across the board - scanner and admin pages that normally load in well under a second took many times longer - and where a database response exceeded the application timeout, the page returned an error rather than loading. What this looked like in practice: **Warehouse floor:** scanner pages \(picking, moving, adjusting inventory\) loaded very slowly; a worker who retried while a page was stuck could receive an error page and have to go back and repeat the action. **Errors were concentrated in a single burst rather than spread across the incident.** Most of them fell within one 15-minute window at the peak of the congestion \(06:15 - 06:30 PDT\). Counting confirmed error pages in our web-server and application logs, the most affected warehouse saw 71 in that window. **The system remained up throughout.** Every submitted transaction either completed correctly or failed cleanly before making any change. During the deepest slowdown window we can show hundreds of transactions completing successfully for the users who continued working. ### What was NOT affected **Data integrity.** The errors occurred at the very start of request processing, before any change was made. No transaction was lost, duplicated, or partially applied. Every fulfillment, inventory move, and shipment posting that completed did so correctly - we verified the integration logs for the incident window. **Order and shipment synchronization to ERPs and marketplaces** completed correctly throughout; postings that queued up during the slowdown were delivered in full during the catch-up \(verified in integration logs - no failed postings\). **All other environments.** Warehouses on our other database instances, and the entire TMS platform, operated normally. **Security and tenancy.** No security boundary was involved at any point. The third-party service in question is a data-integration vendor operating under credentials we issued; the issue was the volume of its reads, not any unauthorized access. ### Timeline \(all times PDT; add 3 hours for ET\) | Time | Event | | --- | --- | | Aug 24, 22:54 - 22:57 | The ETL service restarts ~22 sync jobs against the database within a three-minute window. Each begins re-reading historical change data at full speed. | | Aug 24, 22:54 - 23:50 | Read volume climbs to roughly five times the expected peak, consuming most of the disk bandwidth available to the instance. Overnight warehouse traffic is light, so there is no customer-visible effect yet. | | Aug 25, ~04:30 | US-East warehouse morning shifts begin. Combined demand exceeds the capped network speed; queues start building and the first pages begin rendering slower than normal. | | 05:30 | Automated response-time monitoring alerts as warehouse activity ramps up; customer reports of slowness arrive in the same period. Investigation begins. | | 06:00 - 07:00 | Peak congestion: database connections spike to ~15x normal as requests pile up; the wave of scanner-screen errors occurs \(06:15-06:30\). | | 06:50 | Root cause identified: disk bandwidth over the instance limit; the vendor's replication streams identified as the driver. | | 07:00 | Heaviest internal report queries disabled to free capacity - partial relief. | | 07:20 - 08:30 | ETL service sync jobs are paused in waves in its console and its database sessions terminated; the service automatically reconnects within seconds each time and continues reading. During this period it starts additional jobs. | | 08:35 | The ETL service database credentials are locked and its sessions terminated a final time. | | 08:38 - 08:45 | Database queues drain; page response times return to normal. Customer impact ends. | ### Why resolution took ~3 hours from first reports Three factors extended the timeline. First, the trigger occurred seven hours before symptoms. The ETL service's re-read ran overnight and had already consumed the available bandwidth, but with warehouse activity light at that hour the constraint produced only a slight change in system response times - below our alerting thresholds - so it went undetected. Our automated monitoring did alert once warehouse activity ramped up in the morning, but by then the underlying change was seven hours old and there was no recent deployment or configuration change to point to. Second, the ETL service's replication reads are invisible to standard database query logs - they use a replication protocol rather than queries - so identifying them as the consumer required correlating disk, network, and connection-level evidence. Third, the ETL service is built to survive interruptions: pausing its jobs and terminating its connections both failed as mitigations because it reconnects automatically within seconds, and it restarted additional jobs while we were pausing others. Only revoking its credentials stopped it. ### What we are changing **Tuning WMS response-time alert thresholds.** The condition behind this incident was present for seven hours overnight, but under light load it moved response times too little to cross our alert thresholds - so the first alert came only once warehouse activity ramped up and customers were already affected. We are tuning those thresholds to be sensitive to smaller shifts in WMS response time, including at low load, so events like this are caught and acted on before they reach customers. This includes alerting on the specific leading indicators of this incident - disk bandwidth consumption and disk queue depth. **The database has been migrated to an instance type with substantially more disk bandwidth**, giving significant headroom above peak demand to absorb spikes of this kind. **We are continuing our investigation with the ETL vendor.** We have an open case with them seeking an explanation for the simultaneous job restart, and requiring rate limiting and concurrency caps for re-reads against customer sources. That work is ongoing.