MindTouch incident

CXone Knowledge Management - – Monitoring complete. Status = All Services Running Normally

Major Resolved View vendor source →

MindTouch experienced a major incident on June 26, 2026 affecting Application (General Service) and Search and 1 more component, lasting 2h 18m. The incident has been resolved; the full update timeline is below.

Started
Jun 26, 2026, 04:27 PM UTC
Resolved
Jun 26, 2026, 06:46 PM UTC
Duration
2h 18m
Detected by Pingoru
Jun 26, 2026, 04:27 PM UTC

Affected components

Application (General Service)SearchGenerative SearchIn-Product Contextual HelpEmail ServicesMindTouch Success CenterAnalytics

Update timeline

  1. investigating Jun 26, 2026, 04:27 PM UTC

    CXone Knowledge Management Service Degradation: Sites are slow to load or inaccessible. The CXone Knowledge Management Engineering team is investigating reports of site unavailability.

  2. investigating Jun 26, 2026, 04:45 PM UTC

    CXone Mpower Expert Service Degradation: Sites unavailable. The CXone Knowledge Management Engineering team is investigating reports of site unavailability

  3. identified Jun 26, 2026, 04:55 PM UTC

    CXone Knowledge Management Service Degradation: Sites unavailable. The issue has been identified and a fix is being worked on for deployment.

  4. identified Jun 26, 2026, 05:12 PM UTC

    CXone Knowledge Management Service Degradation: Sites unavailable. The issue has been identified and a fix is being worked on for deployment.

  5. identified Jun 26, 2026, 05:31 PM UTC

    CXone Knowledge Management Service Degradation: Sites unavailable. The issue has been identified and a fix is being worked on for deployment.

  6. monitoring Jun 26, 2026, 05:36 PM UTC

    CXone Knowledge Management - Fix Deployed - All Services Running Normally. The CXone Knowledge Management Engineering team has deployed a fix and all services are running normally. We are currently monitoring sites for deployment stability. Event duration 1h 8m

  7. monitoring Jun 26, 2026, 06:03 PM UTC

    CXone Knowledge Management - Fix Deployed - All Services Running Normally. The CXone Knowledge Management Engineering team has deployed a fix and all services are running normally. We are currently monitoring sites for deployment stability. Event duration 1h 8m

  8. monitoring Jun 26, 2026, 06:27 PM UTC

    CXone Knowledge Management - Fix Deployed - All Services Running Normally. The CXone Knowledge Management Engineering team has deployed a fix and all services are running normally. We are currently monitoring sites for deployment stability. Event duration 1h 8m

  9. resolved Jun 26, 2026, 06:46 PM UTC

    CXone Knowledge Management - Service Disruption Resolved - All Services Running Normally. The CXone Mpower Expert Engineering team has deployed a fix and monitored the deployment to make sure sites are stable. The issue is now resolved at this time. Event duration 1h 8 mins

  10. postmortem Jul 07, 2026, 05:19 PM UTC

    **Impact Start Time \(UTC\) 06/26/2026 04:07 PM UTC** **Impact End Time \(UTC\) 06/26/2026 05:36 PM UTC** **Incident Summary:** On 06/26/2026, some NiCE CXone Mpower customers reported service degradation while accessing and using the CXone Mpower Expert knowledge portal. The recurring service degradation was caused by a Domain Name System \(DNS\) resolution bottleneck within the platform infrastructure during periods of elevated traffic. The issue was mitigated through an emergency change that introduced timeout controls to prevent backend processes from remaining stuck for extended periods, allowing orphaned processes to be released and system resources to be recovered. Service stability was further aided by a reduction in traffic volume. ## Root Cause The recurring service degradation was caused by a DNS resolution bottleneck within the platform infrastructure during periods of elevated traffic. Increased DNS request volume exceeded infrastructure processing limits, resulting in intermittent DNS lookup failure. These failures triggered application retries, increased latency, elevated “5xx” errors, and customer-facing service degradation. The issue was traced to an existing configuration gap that prevented application workloads from utilizing the platform's DNS caching capability. As a result, DNS requests continued to follow the standard resolution path, creating unnecessary network overhead and contributing to DNS request saturation under higher traffic volumes. While this condition had existed for some time without widespread impact, continued platform growth, increased workload density, and infrastructure scaling changes increased DNS traffic and exposed the underlying limitation. Once the DNS processing limits were reached, requests were intermittently dropped, resulting in cascading application failures and customer-facing service disruptions. During the investigation, engineers implemented several mitigations to reduce customer impact while continuing to identify the underlying cause. However, because these actions addressed symptoms rather than the root issue, the condition recurred and resulted in multiple major incidents before a permanent corrective solution was implemented. ## Corrective Actions **Detection:** The failure condition triggered system alarms, prompting engineers to initiate incident response procedures to validate, investigate, and reproduce the issue. While the investigation was underway, some customers reported service degradation when accessing and using the CXone Mpower Expert knowledge portal. **Remediation:** The issue was mitigated through an emergency change that introduced timeout controls to prevent backend processes from remaining stuck for extended periods, allowing orphaned processes to be released and system resources to be recovered. Service stability was further aided by a reduction in traffic volume. Completed on 06/26/2026. **Prevention:** Engineers developed a dashboard to monitor local DNS cache performance, improving visibility into DNS health and enabling earlier detection, investigation, and remediation of potential issues before they impact service availability or customer experience. Completed on 07/01/2026. The Engineering team implemented an emergency change to correctly enable local DNS caching for applicable workloads, reducing DNS traffic, preventing node-level DNS saturation, and restoring the intended DNS architecture. This improved platform stability and DNS performance while reducing the risk of similar incidents. Additional infrastructure optimizations were also implemented to improve operational efficiency and minimize service impact during scaling activities. Completed on 07/02/2026. Engineers will conduct a comprehensive validation of DNS configurations across all workloads to ensure consistent use of node local DNS cache. While the immediate issue has been remediated, a fleet-wide review will identify and correct any remaining configuration gaps, helping to optimize DNS performance, reduce infrastructure load, and prevent similar service degradation events in the future. In addition, the Engineering team will continue extended monitoring of platform health and stability to verify the long-term effectiveness of the corrective actions and proactively identify any emerging issues. An update will be provided by End of Day MT on 07/17/2026. NiCE resiliency teams remain focused on enhancing system reliability, stability, platform-wide resilience, monitoring capabilities, and overall performance. These ongoing efforts are aligned with the Operational Resiliency Excellence initiative, driving continuous improvement, strengthening operational maturity, and delivering measurable business outcomes. ## Incident Timeline \(UTC\) 06/26/2026 04:07 PM \(UTC\) - Engineers received a service degradation alert and immediately initiated incident response procedures to investigate, validate, and reproduce the issue to determine its scope and potential customer impact. 06/26/2026 04:10 PM \(UTC\) - The first customer case was opened, and Tech Support \(TS\) engineers began the troubleshooting investigation. 06/26/2026 04:35 PM \(UTC\) - TS engineers notified the Network Operations Center \(NOC\) engineers about the reported customer impact; a major incident was proposed and confirmed. 06/26/2026 04:55 PM \(UTC\) - Engineers identified a suspected cause and began remediation steps. 06/26/2026 05:36 PM \(UTC\) - The impact was resolved after implementing timeout controls to prevent backend processes from becoming stalled and as traffic levels returned to normal. Following successful validations, the major incident was marked resolved.