MindTouch incident

CXone Knowledge Management - Service Degradation: Sites unavailable

Minor Resolved View vendor source →

MindTouch experienced a minor incident on June 29, 2026 affecting Application (General Service) and Search and 1 more component, lasting 6h 26m. The incident has been resolved; the full update timeline is below.

Started
Jun 29, 2026, 03:25 PM UTC
Resolved
Jun 29, 2026, 09:51 PM UTC
Duration
6h 26m
Detected by Pingoru
Jun 29, 2026, 03:25 PM UTC

Affected components

Application (General Service)SearchIn-Product Contextual HelpEmail ServicesMindTouch Success CenterAnalyticsGeoblocking for Russia

Update timeline

  1. investigating Jun 29, 2026, 03:25 PM UTC

    CXone Knowledge Management Service Degradation. The CXone Knowledge Management Engineering team is investigating reports of degraded performance

  2. investigating Jun 29, 2026, 03:33 PM UTC

    We are continuing to investigate this issue.

  3. investigating Jun 29, 2026, 03:46 PM UTC

    CXone Knowledge Management Service Degradation. The CXone Knowledge Management Engineering team is investigating reports of degraded performance

  4. investigating Jun 29, 2026, 04:02 PM UTC

    CXone Knowledge Management Service Degradation. The CXone Knowledge Management Engineering team is investigating reports of degraded performance

  5. investigating Jun 29, 2026, 04:29 PM UTC

    CXone Knowledge Management Service Degradation. The CXone Knowledge Management Engineering team is investigating reports of degraded performance

  6. monitoring Jun 29, 2026, 04:34 PM UTC

    A fix has been implemented and we are monitoring the results.

  7. resolved Jun 29, 2026, 09:51 PM UTC

    This incident has been resolved.

  8. postmortem Jul 07, 2026, 05:27 PM UTC

    **Impact Start Time \(UTC\) 06/29/2026 01:22 PM UTC** **Impact End Time \(UTC\) 06/29/2026 04:31 PM UTC** **Incident Summary:** On 06/29/2026, some NiCE CXone Mpower customers reported service degradation while accessing and using the CXone Mpower Expert knowledge portal. The recurring service degradation was caused by a Domain Name System \(DNS\) resolution bottleneck within the platform infrastructure during periods of elevated traffic. The impact was resolved after traffic volumes returned to normal levels. ## Root Cause The recurring service degradation was caused by a DNS resolution bottleneck within the platform infrastructure during periods of elevated traffic. Increased DNS request volume exceeded infrastructure processing limits, resulting in intermittent DNS lookup failure. These failures triggered application retries, increased latency, elevated “5xx” errors, and customer-facing service degradation. The issue was traced to an existing configuration gap that prevented application workloads from utilizing the platform's DNS caching capability. As a result, DNS requests continued to follow the standard resolution path, creating unnecessary network overhead and contributing to DNS request saturation under higher traffic volumes. While this condition had existed for some time without widespread impact, continued platform growth, increased workload density, and infrastructure scaling changes increased DNS traffic and exposed the underlying limitation. Once the DNS processing limits were reached, requests were intermittently dropped, resulting in cascading application failures and customer-facing service disruptions. During the investigation, engineers implemented several mitigations to reduce customer impact while continuing to identify the underlying cause. However, because these actions addressed symptoms rather than the root issue, the condition recurred and resulted in multiple major incidents before a permanent corrective solution was implemented. ## Corrective Actions **Detection:** Corrective Actions The failure condition triggered system alarms, prompting engineers to initiate incident response procedures to validate, investigate, and reproduce the issue. While the investigation was underway, some customers reported service degradation when accessing and using the CXone Mpower Expert knowledge portal. **Remediation:** The impact was resolved after traffic volumes returned to normal levels. Completed on 06/29/2026. **Prevention:** Engineers developed a dashboard to monitor local DNS cache performance, improving visibility into DNS health and enabling earlier detection, investigation, and remediation of potential issues before they impact service availability or customer experience. Completed on 07/01/2026. The Engineering team implemented an emergency change to correctly enable local DNS caching for applicable workloads, reducing DNS traffic, preventing node-level DNS saturation, and restoring the intended DNS architecture. This improved platform stability and DNS performance while reducing the risk of similar incidents. Additional infrastructure optimizations were also implemented to improve operational efficiency and minimize service impact during scaling activities. Completed on 07/02/2026. Engineers will conduct a comprehensive validation of DNS configurations across all workloads to ensure consistent use of node local DNS cache. While the immediate issue has been remediated, a fleet-wide review will identify and correct any remaining configuration gaps, helping to optimize DNS performance, reduce infrastructure load, and prevent similar service degradation events in the future. In addition, the Engineering team will continue extended monitoring of platform health and stability to verify the long-term effectiveness of the corrective actions and proactively identify any emerging issues. An update will be provided by End of Day MT on 07/17/2026. NiCE resiliency teams remain focused on enhancing system reliability, stability, platform-wide resilience, monitoring capabilities, and overall performance. These ongoing efforts are aligned with the Operational Resiliency Excellence initiative, driving continuous improvement, strengthening operational maturity, and delivering measurable business outcomes. ## Incident Timeline \(UTC\) 06/29/2026 01:22 PM \(UTC\) - Engineers received a service degradation alert and immediately initiated incident response procedures to investigate, validate, and reproduce the issue to determine its scope and potential customer impact. 06/29/2026 01:30 PM \(UTC\) - The first customer case was reported, and Technical Support \(TS\) engineers began validation and troubleshooting activities. 06/29/2026 03:34 PM \(UTC\) - TS engineers notified the Network Operations Center \(NOC\) engineers about the reported customer impact; a major incident was proposed and confirmed. 06/29/2026 03:39 PM \(UTC\) - Engineers identified a suspected contributing factor and initiated mitigation efforts to stabilize the service. 06/29/2026 04:31 PM \(UTC\) - Service performance recovered as traffic volumes returned to normal levels. Following successful validation of platform