All Systems are Online

INC0120326 CITZ - MCS SILVER - Unexpected reboot for worker node MCS-SILVER-APP-92.DMZ

Resolved

This worker node rebooted itself due to hardware issues. The reboot event caused the node to enter a NotReady state which forced the cluster to evict the existing pods to other nodes. A manual cordon and drain afterward has ensured this node remains empty of workloads while we work with vendor support to troubleshoot the hardware.


Posted: Tue, 18 Aug 2026 05:33:00 +0000

Metrics

Private Cloud - Silver Cluster

Uptime - %

Silver Error Budget

- mins

Private Cloud - Gold Cluster Gold DR

Uptime - %

Gold and Gold DR Error Budget

- mins

Private Cloud - Emerald Cluster

Uptime - %

Emerald Error Budget

- mins

Registry

Uptime - %

Vault

Uptime - %

Artifactory

Uptime - %

Private Cloud - Klab Cluster

Uptime - %

Private Cloud - Clab Cluster

Uptime - %

Overview

Silver
Good Service
Gold
Good Service
GoldDR
Good Service
Emerald
Good Service
Vault
Good Service
Artifactory
Good Service

Recent Events

Last 7 Days

Thursday, August 27th 2026

There are no reported events.

Wednesday, August 26th 2026

There are no reported events.

Tuesday, August 25th 2026

There are no reported events.

Monday, August 24th 2026

CHG0080162 - CITZ MCS GOLD - Apply updates and perform maintenance

» View Event Details | Created Mon, 24 Aug 2026 13:00:00 +0000

Scheduled

We will update some software, firmware, and cluster settings and perform a rolling restart of all nodes in the cluster.

We will apply the latest firmware on the physical servers.

We will upgrade Trident to v26.06.0 and enable concurrency. This has shown great speed improvements to PVC mount/unmount operations in the LABs.

Switch the container runtime from runc to crun, as it is the default for OpenShift now.

When?

Aug 24th to Aug 28th

Trident will be upgraded at 6am on the first day of the change.

At 9am on the first day of the change we will begin restarting the master nodes first, then the infra nodes, and start on the worker nodes last.

Node reboots will continue until 5pm each day, and resume at 9am the next day. Silver will pause over the weekend.

Will there be an impact on the Platform apps?

As nodes are rebooted pods will be rescheduled to another node. We expect the past issues with Trident taking a long time to start up pods will be solved, but will continue to use slow manual node drains and monitoring of Trident health.

Do I need to do anything?

Monitor your apps health.


Posted: Mon, 24 Aug 2026 13:00:00 +0000

Sunday, August 23rd 2026

There are no reported events.

Saturday, August 22nd 2026

There are no reported events.

Friday, August 21st 2026

There are no reported events.

Upcoming Events

Next 30 Days

Monday, August 31st 2026

CHG0080163 - CITZ MCS GOLDDR - Apply updates and perform maintenance

» Mon, 31 Aug 2026 13:00:00 +0000

Scheduled

We will update some software, firmware, and cluster settings and perform a rolling restart of all nodes in the cluster.

We will apply the latest firmware on the physical servers.

We will upgrade Trident to v26.06.0 and enable concurrency. This has shown great speed improvements to PVC mount/unmount operations in the LABs.

Switch the container runtime from runc to crun, as it is the default for OpenShift now.

When?

Aug 31st to Sept 4th

Trident will be upgraded at 6am on the first day of the change.

At 9am on the first day of the change we will begin restarting the master nodes first, then the infra nodes, and start on the worker nodes last.

Node reboots will continue until 5pm each day, and resume at 9am the next day. Silver will pause over the weekend.

Will there be an impact on the Platform apps?

As nodes are rebooted pods will be rescheduled to another node. We expect the past issues with Trident taking a long time to start up pods will be solved, but will continue to use slow manual node drains and monitoring of Trident health.

Do I need to do anything?

Monitor your apps health.


Posted: Mon, 31 Aug 2026 13:00:00 +0000

CHG0080164 - CITZ MCS EMERALD - Apply updates and perform maintenance

» Mon, 31 Aug 2026 13:00:00 +0000

Scheduled

We will update some software, firmware, and cluster settings and perform a rolling restart of all nodes in the cluster.

We will apply the latest firmware on the physical servers.

We will upgrade Trident to v26.06.0 and enable concurrency. This has shown great speed improvements to PVC mount/unmount operations in the LABs.

Switch to CGroups v2 to unblock future OpenShift upgrades.

Switch the container runtime from runc to crun, as it is the default for OpenShift now.

When?

Aug 31st to Sept 4th

Trident will be upgraded at 6am on the first day of the change.

At 9am on the first day of the change we will begin restarting the master nodes first, then the infra nodes, and start on the worker nodes last.

Node reboots will continue until 5pm each day, and resume at 9am the next day. Silver will pause over the weekend.

Will there be an impact on the Platform apps?

As nodes are rebooted pods will be rescheduled to another node. We expect the past issues with Trident taking a long time to start up pods will be solved, but will continue to use slow manual node drains and monitoring of Trident health.

Do I need to do anything?

Monitor your apps health.


Posted: Mon, 31 Aug 2026 13:00:00 +0000

Subscribe to Updates