Identity Services Outage affecting customers
Incident window: June 23, 2023 13:16 UTC to 14:00 UTCImpacted Cloud services:
Foundry Identity Service
Foundry Identity Services were running unhealthy in US-EAST Virgina region.
Impact Level : high
Identity Services is being investigated.
[2023-06-23 13:16 UTC] Foundry Identity Services were running unhealthy. This could have impacted multiple customers.
[2023-06-23 14:00 UTC] The services are restored back and it is running healthy.
[2023-06-23 16:21 UTC] We have captured all the logs and required information, our internal teams are investigating the issue.
[2023-06-26 17:21 UTC] The services have been continuing to perform within normal parameters. We are continuing our investigation to identify the root cause.
[2023-06-27 18:45 UTC] The services have been continuing to perform within normal parameters. We are continuing our investigation to identify the root cause.
[2023-06-28 14:00 UTC] The services have been continuing to perform within normal parameters. We are continuing our investigation to identify the root cause.
[2023-06-29 12:45 UTC] The services have been continuing to perform within normal parameters. We are continuing our investigation to identify the root cause.
[2023-06-30 10:00 UTC] The services have been continuing to perform within normal parameters. We are continuing our investigation to identify the root cause.
[2023-07-01 15:01 UTC] The services have been continuing to perform within normal parameters. We are continuing our investigation to identify the root cause.
[2023-07-02 16:01 UTC] The services have been continuing to perform within normal parameters. We are continuing our investigation to identify the root cause.
[2023-07-03 10:01 UTC] The services have been continuing to perform within normal parameters. We are continuing our investigation to identify the root cause.
Root Cause Analysis: During this incident It was observed there were evidences of failed communication between the Authentication VM hosted on US EAST region and one of the internal data source with IOExceptions followed by connection reset impacting the customers traffic with an authentication failure. IOExceptions were sourced from an unhealthy state of the connection objects in the VM.
Remediation: As this is a very very rare scenario to track down we have positioned additional automated observability with different anomaly conditions and are being observed spontaneously. In order to prevent such cases in future VoltMX Engineering is putting additional capabilities within is the product an improved visibility into the object’s health for better management of these objects in the VM.