Open role
Senior Staff Site Reliability Engineer – Compute Platform
NVIDIA
India, BengaluruPosted Sep 17, 2026 · 3h ago
Full-timeStaffOn-siteConsumer Tech
About this role
NVIDIA is seeking a Senior Staff Site Reliability Engineer to build and operate scalable compute platforms for global engineering workloads. This role involves managing Kubernetes, KubeVirt, and bare-metal infrastructure, focusing on automation, observability, and AI-enabled operations to solve complex infrastructure challenges.
What we are looking for
6- Build and operate large-scale Kubernetes, KubeVirt, Linux, and bare-metal compute platforms
- Lead bare-metal provisioning and lifecycle management in data centers
- Develop automation and observability solutions using Python or Go
- Define and operate SLOs, SLIs, and incident-response practices
- Partner with infrastructure and application teams for global platform initiatives
- Participate in an on-call rotation
Skills mentioned
5PythonDockerKubernetesLinuxTerraform
