LLM Gateway service interruption
Timeline · 1 update
- resolved Oct 08, 2026, 07:42 AM UTC
Status: Resolved LLM Gateway is not servicing incoming requests due to internal misconfiguration.
LangChain had 35 outages in the last 2 years totaling 5h 3m of downtime — averaging 1.4 incidents per month.
There were 35 LangChain outages since November 14, 2025 totaling 5h 3m of downtime. Each is summarised below — incident details, duration, and resolution information.
Status: Resolved LLM Gateway is not servicing incoming requests due to internal misconfiguration.
Status: Investigating Creation of new LangSmith deployment dedicated instances are failing during database provisioning because of an issue with an upstream cloud provider. Only GCP regions are affected. Affected components LangSmith Deployments Control Plane (Partial outage)
Status: Investigating Creating new deployments and new revisions for dedicated instances is failing for LangSmith Deployment in all GCP regions due to an issue with an upstream cloud provider. Existing deployments continue to run and serve traffic. Affected components LangSmith Deployments Control Plane (Partial outage)
Status: Identified Issue has been identified, attempted mitigation is being deployed. Affected components LangSmith Deployments Control Plane (Partial outage)
Status: Resolved This is now resolved. Affected components LangSmith Deployments Control Plane (Operational)
Status: Investigating This is impacting the creation of new agents with a computer per thread configuration Affected components LangSmith Fleet (Degraded performance)
Status: Identified We have identified the cause and are working on resolving the issue. Affected components LangSmith Fleet (Degraded performance)
Status: Monitoring The fix has been deployed. All systems are operational Affected components LangSmith Fleet (Operational)
Status: Resolved All systems are operational Affected components LangSmith Fleet (Operational)
Status: Investigating This affects only newly created agents with a shared computer. Affected components LangSmith Fleet (Partial outage)
Status: Identified We have identified the root cause of this incident and are actively working on a fix. The issue originates from a recent change to how sandbox slugs are generated in our backend systems, which now derive slugs directly rather than referencing sandbox IDs. This has introduced a mismatch between the values returned by the backend and those the frontend expects. We will share a further update as remediation progresses. Affected components LangSmith Fleet (Partial outage)
Status: Monitoring The fix has been deployed. All systems are operational Affected components LangSmith Fleet (Operational)
Status: Resolved All systems are operational Affected components LangSmith Fleet (Operational)
Status: Investigating We are investigating an elevated error rate for bulk exports executions. Bulk exports are currently delayed. New bulk exports can still be created. Affected components LangSmith Bulk Exports (Partial outage)
Status: Monitoring We have mitigated the issue and have seen recovery. We're continuing to monitor. Affected components LangSmith Bulk Exports (Operational)
Status: Resolved This incident has been resolved. Affected components LangSmith Bulk Exports (Operational)
Status: Investigating We are currently investigating an issue where some trace loads are timing out
Status: Investigating Currently dealing with some heavy queries causing issues fetching traces. We are scaling out our system which should remediate.
Status: Investigating This is primarily affecting run querying
Status: Monitoring We've identified the query causing problems to our systems and have pushed a mitigation. All systems look recovered but we are monitoring for any other problems.
Status: Resolved A pathological query was impacting other query services. We've added protections against the pathological case and scaled up our query services. All services are now functioning as normal
Status: Investigating Builds for new revisions of LangSmith Deployments are not being processed. Affected components LangSmith Deployments Control Plane (Partial outage)
Status: Monitoring Builds are being processed. Validating recovery. Affected components LangSmith Deployments Control Plane (Partial outage)
Status: Identified GCP is having an incident in us-west-1. Deployments may experience issues deploying new revisions. Affected components LangSmith Deployments Control Plane (Degraded performance)
Status: Monitoring New revisions are being processed normally. Validating recovery. Affected components LangSmith Deployments Control Plane (Degraded performance)
Status: Resolved New revisions are deploying correctly. Affected components LangSmith Deployments Control Plane (Operational)
Status: Identified Traces including large payloads might currently fail to be ingested or can be delayed. We are working on a fix.
Status: Identified Updated affected regions
Status: Monitoring The service is temporarily patched. Large payloads are now ingested normally.
Status: Resolved The full fix took an extraordinarily long time to release due to an outage at our CI provider. In the meantime, the mitigation we had deployed mitigated the problem.
Status: Investigating We're investigating an issue where agents and workflows using Anthropic Claude models with custom tools may fail to complete runs. Affected requests return a tool input validation error. We'll share more information as we have it. Affected components LangSmith Fleet (Degraded performance)
Status: Identified We're seeing elevated error rates when using Fleet with Claude models and certain MCP tools Affected components LangSmith Fleet (Degraded performance)
Status: Monitoring We're seeing elevated error rates when using Fleet with Claude models and certain MCP tools. Affected components LangSmith Fleet (Degraded performance)
Status: Resolved Resolved. Fix is now deployed that skips an unsupported tool schema and logs a warning instead of failing the request. Affected components LangSmith Fleet (Operational)
Status: Investigating We are aware of an issue causing high latency and intermittent HTTP 500 errors for analytical queries (Stats and Dashboards) in LangSmith. Our team is actively investigating and working on a resolution.
Status: Investigating We are aware of an issue causing high latency and intermittent HTTP 500 errors for analytical queries (Stats and Dashboards) in LangSmith. Our team is actively investigating and working on a resolution.
Status: Monitoring We have identified the root cause, mitigated the issue, and are actively monitoring the service to ensure continued stability.
Status: Resolved The issue has been resolved, and service has returned to normal. We identified the root cause, implemented a mitigation, and have observed stable system performance following the change. We will continue to monitor the service closely to ensure ongoing stability. Thank you for your patience while we worked to resolve this issue.
Status: Identified New revisions for Github backed projects are failing. This is due to an ongoing Github incident. Affected components LangSmith Deployments Control Plane (Degraded performance)
Status: Resolved The Github incident has resolved. New revisions are being processed correctly. Affected components LangSmith Deployments Control Plane (Operational)
Status: Investigating We are seeing some failures in trace ingestion. Ingestion is currently delayed. Affected components LangSmith Run Ingestion (Partial outage)
Status: Resolved All systems seem fully recovered. Ingestion was down for approximately 30 minutes from 5pm EST to 530 EST. Traces from that period are still being being processed but there was no data loss. Affected components LangSmith Run Ingestion (Operational)
Status: Investigating We're investigating an issue affecting Fleet where users are unable to interact with their agents. Our team is actively working on a fix and will provide updates as we learn more. Affected components LangSmith Fleet (Partial outage)
Status: Monitoring We've deployed a fix and are seeing error rates return to normal. We're monitoring Fleet closely to confirm full recovery. Affected components LangSmith Fleet (Degraded performance)
Status: Resolved The issue affecting Fleet has been resolved, and users can interact with their agents normally again. Thanks for your patience. Affected components LangSmith Fleet (Operational)
Status: Investigating We are experiencing a performance degradation in one of our legacy storage clusters after a recent version update. This could lead to timeouts when querying trace information. Work is underway to revert the update. Ingestion of trace data is not impacted. Affected components LangSmith Application (Degraded performance) LangSmith API (Degraded performance)
Status: Monitoring The rollback of the legacy storage cluster is mostly complete, and request latencies have largely normalized. Affected components LangSmith Application (Degraded performance) LangSmith API (Degraded performance)
Status: Resolved The rollback is complete, and our latency readings are back to nominal levels. Affected components LangSmith API (Operational) LangSmith Application (Operational)
Status: Investigating We are investigating elevated error rates in the LangSmith API. Affected components LangSmith Application (Partial outage) LangSmith API (Partial outage)
Status: Monitoring We've mitigated the issue and are seeing initial recovery and are continuing to monitor. Affected components LangSmith Application (Operational) LangSmith API (Operational)
Status: Resolved This incident has been resolved. Affected components LangSmith Application (Operational) LangSmith API (Operational)
Status: Identified We have identified an issue causing delays in billable usage tracking for LangSmith Deployments. Affected components LangSmith Billing (Degraded performance)
Status: Monitoring We have a fix in place and are processing the backlog of usage. Affected components LangSmith Billing (Degraded performance)
Status: Resolved The backlog has completed processing. This incident is resolved. Affected components LangSmith Billing (Operational)
Status: Investigating We're currently investigating an issue causing a number of bulk exports to become stuck. Affected components LangSmith API (Partial outage)
Status: Monitoring A fix has been successfully deployed and bulk exports are succeeding again. The team continue to monitor for any further issues. Affected components LangSmith API (Operational)
Status: Resolved No further issues have been observed and this incident is now considered resolved. Affected components LangSmith API (Operational)
Status: Investigating All sandbox create calls error Affected components LangSmith Sandboxes (Partial outage)
Status: Identified Fix is rolling out Affected components LangSmith Sandboxes (Partial outage)
Status: Monitoring Creates succeeding again Affected components LangSmith Sandboxes (Operational)
Status: Resolved Resolved Affected components LangSmith Sandboxes (Operational)
Status: Investigating We're currently investigating an issue causing some bulk exports to fail or become stuck. Affected components LangSmith API (Partial outage)
Status: Identified The team have identified the issue and a fix is being deployed. Affected components LangSmith API (Partial outage)
Status: Monitoring The fix has been successfully deployed and bulk exports are succeeding again. The team continue to monitor for any further issues. Affected components LangSmith API (Operational)
Status: Investigating Unfortunately we continue to see a number of bulk export jobs failing. The team are investigating the cause. Affected components LangSmith API (Partial outage)
Status: Resolved This incident has been resolved. Affected components LangSmith API (Operational)
Status: Investigating We are investigating an increase in backend latency in the LangSmith service. Affected components LangSmith Run Ingestion (Degraded performance) LangSmith Application (Degraded performance) LangSmith API (Degraded performance)
Status: Identified We identified a database modification that came through our regular release pipeline and caused backend contention, resulting in overall slowness in our API and WebUI. Affected components LangSmith Application (Degraded performance) LangSmith API (Degraded performance) LangSmith Run Ingestion (Degraded performance)
Status: Resolved The issue is now fully resolved. All services are restored and fully operational. Affected components LangSmith Run Ingestion (Operational) LangSmith Application (Operational) LangSmith API (Operational)
Status: Investigating We're currently investigating an incident affecting trace visualization in the LangSmith UI. Traces may not render correctly or may fail to load. Affected components LangSmith Application (Partial outage)
Status: Identified We have identified the root cause and are working on a fix. Affected components LangSmith Application (Partial outage)
Status: Monitoring We have deployed a fix and are continuing to monitor the situation closely. Trace ingestion was not impacted. We apologize for the inconvenience. Affected components LangSmith Application (Partial outage)
Status: Resolved The fix has been properly released. Thank you for your patience and sorry for the inconvenience. Affected components LangSmith Application (Operational)
Status: Investigating We're investigating an issue where the annotation queue page may load very slowly or even freeze the UI after clicking the View Run button. Affected components LangSmith Application (Degraded performance)
Status: Identified The cause of the issue has been identified and the team is currently working on a fix. Affected components LangSmith Application (Degraded performance)
Status: Resolved This incident has been resolved. Annotation queues are once again fully accessible in the US LangSmith UI. Please let us know at [email protected] if you're still experiencing issues. Affected components LangSmith Application (Operational)
Status: Investigating We're currently investigating an issue causing deployment failures. Affected components LangSmith Deployments Data Plane (Partial outage)
Status: Identified The issue is impacting the LangSmith Deployments Control Plane. A fix is being deployed now. Affected components LangSmith Deployments Control Plane (Partial outage) LangSmith Deployments Data Plane (Operational)
Status: Monitoring A fix for this issue has been deployed. The team continue to monitor system health. Affected components LangSmith Deployments Control Plane (Operational) LangSmith Deployments Data Plane (Operational)
Status: Resolved The fix was successful and this issue is now considered resolved. Affected components LangSmith Deployments Control Plane (Operational) LangSmith Deployments Data Plane (Operational)
Status: Identified We've identified a capacity issue causing stuck and delayed bulk exports. We're working on increasing capacity to mitigate. Affected components LangSmith API (Degraded performance)
Status: Identified We've scaled up the relevant systems and are observing recovery as the queue drains. Affected components LangSmith API (Degraded performance)
Status: Monitoring The backlog is continuing to process, and we are continuing to monitor. Affected components LangSmith API (Degraded performance)
Status: Resolved This incident has been resolved. Affected components LangSmith API (Operational)
Status: Monitoring We have just experienced a ~18 minute period between 3:24-3:42PM PDT of elevated 5XX error rates impacting all endpoints including trace ingestion. Our systems have recovered and we are monitoring closely. Affected components LangSmith Application (Partial outage) LangSmith Deployments Control Plane (Partial outage) LangSmith Billing (Partial outage) LangSmith PromptHub (Partial outage) LangSmith API (Partial outage) LangSmith Run Ingestion (Partial outage)
Status: Monitoring Systems are operational, and we are monitoring closely. Affected components LangSmith Deployments Control Plane (Operational) LangSmith Billing (Operational) LangSmith PromptHub (Operational) LangSmith API (Operational) LangSmith Run Ingestion (Operational) LangSmith Application (Operational)
Status: Resolved All services were restored and fully operational on February 21, 2026, at 23:42 UTC. Affected components LangSmith Deployments Control Plane (Operational) LangSmith Billing (Operational) LangSmith PromptHub (Operational) LangSmith API (Operational) LangSmith Run Ingestion (Operational) LangSmith Application (Operational)