menu

Feed available - Subscribe to our feed to stay up to date on upcoming maintenance and incidents.

Identity Services Outage affecting customers

Incident window: June 23, 2023 13:16 UTC to 14:00 UTC

Impacted Cloud services:

Impact Level : high

Identity Services is being investigated.

[2023-06-23 13:16 UTC] Foundry Identity Services were running unhealthy. This could have impacted multiple customers.

[2023-06-23 14:00 UTC] The services are restored back and it is running healthy.

[2023-06-23 16:21 UTC] We have captured all the logs and required information, our internal teams are investigating the issue.

[2023-06-26 17:21 UTC] The services have been continuing to perform within normal parameters. We are continuing our investigation to identify the root cause.

[2023-06-27 18:45 UTC] The services have been continuing to perform within normal parameters. We are continuing our investigation to identify the root cause.

[2023-06-28 14:00 UTC] The services have been continuing to perform within normal parameters. We are continuing our investigation to identify the root cause.

[2023-06-29 12:45 UTC] The services have been continuing to perform within normal parameters. We are continuing our investigation to identify the root cause.

[2023-06-30 10:00 UTC] The services have been continuing to perform within normal parameters. We are continuing our investigation to identify the root cause.

[2023-07-01 15:01 UTC] The services have been continuing to perform within normal parameters. We are continuing our investigation to identify the root cause.

[2023-07-02 16:01 UTC] The services have been continuing to perform within normal parameters. We are continuing our investigation to identify the root cause.

[2023-07-03 10:01 UTC] The services have been continuing to perform within normal parameters. We are continuing our investigation to identify the root cause.

Root Cause Analysis: During this incident It was observed there were evidences of failed communication between the Authentication VM hosted on US EAST region and one of the internal data source with IOExceptions followed by connection reset impacting the customers traffic with an authentication failure. IOExceptions were sourced from an unhealthy state of the connection objects in the VM.

Remediation: As this is a very very rare scenario to track down we have positioned additional automated observability with different anomaly conditions and are being observed spontaneously. In order to prevent such cases in future VoltMX Engineering is putting additional capabilities within is the product an improved visibility into the object’s health for better management of these objects in the VM.