AMD GPU Plugin
When you enable the AMD GPU Plugin cluster add-on, you can pass the following key/value pairs as arguments.
Note that to ensure that workloads running on AMD GPU worker nodes are not interrupted unexpectedly, we recommend that you choose the version of the AMD GPU Plugin add-on to deploy, rather than specifying that you want Oracle to update the add-on automatically.
| Key (API and CLI) | Key's Display Name (Console) | Description | Required/Optional | Default Value | Example Value |
|---|---|---|---|---|---|
affinity |
affinity |
A group of affinity scheduling rules. JSON format in plain text or Base64 encoded. Not used by:
Possible equivalents:
|
Optional | null | null |
nodeSelectors |
node selectors |
You can use node selectors and node labels to control the worker nodes on which add-on pods run. For a pod to run on a node, the pod's node selector must have the same key/value as the node's label. Set JSON format in plain text or Base64 encoded. Not used by:
Possible equivalents:
|
Optional | null | {"foo":"bar", "foo2": "bar2"}The pod will only run on nodes that have the |
numOfReplicas |
numOfReplicas | The number of replicas of the add-on deployment. Not used by:
Possible equivalents:
|
Required | 1Creates one replica of the add-on deployment per cluster. |
2Creates two replicas of the add-on deployment per cluster. |
rollingUpdate |
rollingUpdate |
Controls the desired behavior of rolling update by maxSurge and maxUnavailable. JSON format in plain text or Base64 encoded. Not used by:
Possible equivalents:
|
Optional | null | null |
tolerations |
tolerations |
You can use taints and tolerations to control the worker nodes on which add-on pods run. For a pod to run on a node that has a taint, the pod must have a corresponding toleration. Set JSON format in plain text or Base64 encoded. Possible equivalents:
|
Optional | null | [{"key":"tolerationKeyFoo", "value":"tolerationValBar", "effect":"noSchedule", "operator":"exists"}]Only pods that have this toleration can run on worker nodes that have the |
topologySpreadConstraints |
topologySpreadConstraints |
How to spread matching pods among the given topology. JSON format in plain text or Base64 encoded. Not used by:
|
Optional | null | null |
| Key (API and CLI) | Key's Display Name (Console) | Description | Required/Optional | Default Value | Example Value |
|---|---|---|---|---|---|
amd-gpu-device-plugin.ContainerResources |
amd-gpu-device-plugin container resources |
You can specify the resource quantities that the add-on containers request, and set resource usage limits that the add-on containers cannot exceed. JSON format in plain text or Base64 encoded. |
Optional | null |
{"limits": {"cpu": "500m", "memory": "200Mi" }, "requests": {"cpu": "100m", "memory": "100Mi"}}
Create add-on containers that request 100 milllicores of CPU, and 100 mebibytes of memory. Limit add-on containers to 500 milllicores of CPU, and 200 mebibytes of memory. |
pulse
|
Enable health checks |
Time interval in seconds for the plugin to update the kubelet with device health status. Set to |
Optional |
0
|
The following table describes considerations for configuring this cluster add-on in large clusters.
| Argument Name | Roles wrt number of Nodes | Description | As cluster size increases | Risks | Recommendation |
|---|---|---|---|---|---|
amd-gpu-device-plugin.ContainerResources
|
Y |
Defines CPU and memory requests and limits for the AMD GPU Plugin containers. |
The plugin runs on every GPU node. Total resource usage grows linearly with the number of GPU nodes. The following activities increase:
|
If the resources are undersized:
|
Increase the memory and CPU allocated on each node based on the following factors:
Monitor kubelet logs and the CPU and memory usage of the plugin. |
nodeSelectors/affinity
|
Y |
Controls which nodes the plugin runs on by using labels. |
Ensures that the plugin runs only on AMD GPU nodes. Prevents unnecessary scheduling on CPU-only nodes. |
If no selector is specified, the plugin runs on all nodes and wastes resources. If the wrong selector is specified, the plugin does not run on GPU nodes and the GPUs are not visible to Kubernetes. Even a small misconfiguration affecting 1–2% of nodes can result in hundreds of incorrectly configured nodes, inconsistent GPU availability, and unpredictable scheduling failures. |
Use strict node labeling. |
tolerations
|
Y |
Controls whether the plugin can run on tainted GPU nodes. |
If not configured correctly, the AMD GPU Plugin DaemonSet might fail to schedule on GPU nodes with taints, preventing GPU resources from being discovered and making GPUs unavailable for workload scheduling. |
If tolerations are misconfigured, the plugin cannot run on GPU nodes and the GPUs become unusable. |
Ensure that tolerations match the taints applied to GPU nodes. |
pulse
|
Y |
Controls how frequently the plugin reports GPU health information to the kubelet. |
A higher reporting frequency provides better GPU health visibility but generates more kubelet and API traffic. A lower reporting frequency reduces overhead but delays failure detection. |
If updates are too frequent, CPU overhead on each node and kubelet load increase. If updates are too infrequent or disabled, GPU health information can become stale and the scheduler might assign workloads to unhealthy GPUs. |
Use a moderate interval that balances health visibility against overhead. Avoid overly aggressive polling. |
rollingUpdate
|
N |
Controls update behavior. |
In large clusters, many nodes might be updated simultaneously, affecting GPU availability during upgrades. |
|
Use controlled rolling updates to avoid widespread disruption. |
numOfReplicas
|
N |
This argument is not used by the AMD GPU Plugin. |
Not applicable | Not applicable | Not applicable |