Operate
Day-Two Operations
Governed runbooks for the ongoing work of running a GPU platform — upgrades, patching, rotation, scaling, and disaster recovery. Every procedure has pre-checks, validation, and a rollback.
high risk
GPU driver upgrade
Upgrade the NVIDIA/AMD driver on GPU nodes, one at a time, via the GPU Operator, validating before moving on.
Open runbook high riskGPU Operator upgrade
Upgrade the GPU Operator with a reviewed Helm diff and staged rollout.
Open runbook high riskKubernetes version upgrade
Upgrade the control plane, then GPU node pools, one minor version at a time, revalidating GPU enablement after.
Open runbook