NVIDIA GPU Plugin
When you enable the NVIDIA GPU Plugin cluster add-on, you can pass the following key/value pairs as arguments.
Note that to ensure that workloads running on NVIDIA GPU worker nodes are not interrupted unexpectedly, we recommend that you choose the version of the NVIDIA GPU Plugin add-on to deploy, rather than specifying that you want Oracle to update the add-on automatically.
| Key (API and CLI) | Key's Display Name (Console) | Description | Required/Optional | Default Value | Example Value |
|---|---|---|---|---|---|
affinity |
affinity |
A group of affinity scheduling rules. JSON format in plain text or Base64 encoded. Not used by:
Possible equivalents:
|
Optional | null | null |
nodeSelectors |
node selectors |
You can use node selectors and node labels to control the worker nodes on which add-on pods run. For a pod to run on a node, the pod's node selector must have the same key/value as the node's label. Set JSON format in plain text or Base64 encoded. Not used by:
Possible equivalents:
|
Optional | null | {"foo":"bar", "foo2": "bar2"}The pod will only run on nodes that have the |
numOfReplicas |
numOfReplicas | The number of replicas of the add-on deployment. Not used by:
Possible equivalents:
|
Required | 1Creates one replica of the add-on deployment per cluster. |
2Creates two replicas of the add-on deployment per cluster. |
rollingUpdate |
rollingUpdate |
Controls the desired behavior of rolling update by maxSurge and maxUnavailable. JSON format in plain text or Base64 encoded. Not used by:
Possible equivalents:
|
Optional | null | null |
tolerations |
tolerations |
You can use taints and tolerations to control the worker nodes on which add-on pods run. For a pod to run on a node that has a taint, the pod must have a corresponding toleration. Set JSON format in plain text or Base64 encoded. Possible equivalents:
|
Optional | null | [{"key":"tolerationKeyFoo", "value":"tolerationValBar", "effect":"noSchedule", "operator":"exists"}]Only pods that have this toleration can run on worker nodes that have the |
topologySpreadConstraints |
topologySpreadConstraints |
How to spread matching pods among the given topology. JSON format in plain text or Base64 encoded. Not used by:
|
Optional | null | null |
| Key (API and CLI) | Key's Display Name (Console) | Description | Required/Optional | Default Value | Example Value |
|---|---|---|---|---|---|
deviceIdStrategy
|
Device ID Strategy |
Which strategy to use for passing device IDs to the underlying runtime. One of:
|
Optional |
uuid
|
|
deviceListStrategy
|
Device List Strategy |
Which strategy to use for passing the device list to the underlying runtime. Supported values:
Multiple values are supported, in a comma-separated list. |
Optional |
envvar
|
|
driverRoot
|
Driver Root | The root path for the NVIDIA driver installation. | Optional |
/
|
|
failOnInitError
|
FailOnInitError |
Whether to fail the plugin if an error is encountered during initialization. When set to |
Optional |
true
|
|
migStrategy
|
MIG Strategy |
Which strategy to use for exposing MIG (Multi-Instance GPU) devices on GPUs that support it. One of:
|
Optional |
none
|
|
nvidia-gpu-device-plugin.ContainerResources
|
nvidia-gpu-device-plugin container resources |
You can specify the resource quantities that the add-on containers request, and set resource usage limits that the add-on containers cannot exceed. JSON format in plain text or Base64 encoded. |
Optional | null |
{"limits": {"cpu": "500m", "memory": "200Mi" }, "requests": {"cpu": "100m", "memory": "100Mi"}}
Create add-on containers that request 100 milllicores of CPU, and 100 mebibytes of memory. Limit add-on containers to 500 milllicores of CPU, and 200 mebibytes of memory. |
passDeviceSpecs
|
Pass Device Specs | Whether to pass the paths and desired device node permissions for any NVIDIA devices being allocated to the container. | Optional |
false
|
|
useConfigFile
|
Use Config File from ConfigMap |
Whether to use a configuration file to configure the Nvidia Device Plugin for Kubernetes. The configuration file is derived from a ConfigMap. If set to The ConfigMap is referenced by the |
Optional |
false
|
Example of nvidia-device-plugin-config ConfigMap:
apiVersion: v1
kind: ConfigMap
metadata:
name: nvidia-device-plugin-config
namespace: kube-system
data:
config.yaml: |
version: v1
flags:
migStrategy: "none"
failOnInitError: true
nvidiaDriverRoot: "/"
plugin:
passDeviceSpecs: false
deviceListStrategy: envvar
deviceIDStrategy: uuidThe following table describes considerations for configuring this cluster add-on in large clusters.
| Argument Name | Roles wrt number of Nodes | Description | As cluster size increases | Risks | Recommendation |
|---|---|---|---|---|---|
nvidia-gpu-device-plugin.ContainerResources
|
Y |
Defines CPU and memory requests and limits for NVIDIA GPU Plugin pods. |
The plugin runs on every GPU node as a DaemonSet, so total resource usage increases with the number of GPU nodes. Additional GPU nodes increase:
|
If the resources are undersized:
|
|
affinity
/
nodeSelectors
/
tolerations
|
Y |
Control the nodes on which NVIDIA GPU Plugin pods run. |
|
If these arguments are misconfigured:
|
|
topologySpreadConstraints
|
N |
Controls how NVIDIA GPU Plugin pods are distributed across nodes and availability domains. |
Topology spread constraints help maintain an even distribution in multi-domain clusters and clusters with heterogeneous GPU node pools. |
If topology spread constraints are not configured:
|
Use topology spread constraints for large GPU clusters that span multiple availability domains or contain heterogeneous GPU node pools. |
numOfReplicas
|
N |
Controls the number of replicas for components that support replica-based scaling. |
The NVIDIA GPU Plugin runs as a DaemonSet with one pod on each eligible GPU node. It is not scaled by changing a replica count. |
Not applicable | Not applicable |
rollingUpdate
|
N |
Controls how NVIDIA GPU Plugin pods are updated. |
In large clusters, an update can affect many GPU nodes and temporarily reduce GPU availability. |
|
|
failOnInitError
|
Y |
Controls the plugin behavior when initialization fails. |
Initialization failures become more likely to occur somewhere in the cluster as the number of GPU nodes increases. |
|
|
deviceListStrategy
/
deviceIdStrategy
|
N |
Control how GPU devices are identified and exposed to containers. |
The selected strategies affect container-runtime integration and GPU-allocation efficiency across GPU nodes. |
If configured incorrectly, GPU devices might not be identified or injected consistently into containers. This might result in GPU workload failures, incorrect device assignments, or scheduling issues across GPU nodes. |
Configure deviceListStrategy and deviceIdStrategy to match your container runtime and GPU environment, and use the default uuid device identification strategy unless a validated use case requires otherwise. |
migStrategy
|
N |
Controls how NVIDIA Multi-Instance GPU resources are exposed to Kubernetes. |
Multi-Instance GPU can improve GPU sharing and utilization as the number of GPU workloads increases. |
If this argument is misconfigured:
|
Use Multi-Instance GPU only when GPU partitioning is required. |