Open role
Senior Site Reliability Engineer, DGX Cloud
NVIDIA
4 LocationsPosted Sep 16, 2026 · Yesterday$168k – $334k
Full-time$168k – $334kSeniorHybridSaaS
About this role
NVIDIA is seeking a Senior Site Reliability Engineer to join their DGX Cloud team. You will be responsible for building, implementing, and supporting large-scale Kubernetes clusters for AI researchers and enterprise clients on major cloud providers. This role offers the opportunity to work at the forefront of AI and cloud computing, optimizing performance and ensuring the reliability of cutting-edge infrastructure.
What we are looking for
6- Build and support large-scale Kubernetes clusters for AI workloads
- Define and monitor SLOs/SLIs for system health
- Optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds
- Lead incident response and root-cause analysis
- Automate systems for sustainable scaling and reliability
- 8+ years of experience operating production services
Skills mentioned
7PythonAWSAzureGCPKubernetesLinuxTerraform
