OCI Kubernetes Engine (OKE) support for GPU host repair
- Services: Kubernetes Engine
- Release Date: October 02, 2026
You can now use OKE to coordinate the repair of eligible managed bare metal GPU worker nodes in response to specified hardware-related conditions.
The optional GPU Health Monitoring add-on monitors supported GPU health conditions on eligible managed GPU worker nodes in enhanced clusters. You can use the health signals that the add-on reports when configuring GPU instance recovery.
When you or the GPU Health Monitoring add-on determine a node is eligible for repair, OKE coordinates the repair workflow. The workflow first makes the node unavailable for new workloads and attempts to drain workloads from the node before performing the applicable Compute action.
GPU host repair supports the following repair types:
- User-initiated repair: Use this repair type when you identify a condition that requires a GPU worker node to be replaced. You select an eligible node or nodes in need of repair, and OKE coordinates draining the node, reporting the condition to Compute, terminating the affected instance, and restoring managed node pool capacity.
- GPU instance recovery: Use this repair type when GPU health monitoring reports an eligible critical GPU health signal. You select an eligible node or nodes, and when an eligible critical GPU health signal is reported, OKE coordinates draining the node and rebooting the affected instance.
For more information, see Reparing GPU Hosts.