Open role
Senior Staff Site Reliability Engineer
NVIDIA
India, BengaluruPosted Aug 20, 2026 · 3h ago
Full-timeStaffOn-siteEnterprise Software
About this role
NVIDIA is seeking a Senior Staff Site Reliability Engineer to lead the development of the runtime foundation for their enterprise AI platforms. You will define the architecture and technical roadmap for scalable AI runtime systems, design Kubernetes-based infrastructure, and build automation for provisioning, scaling, and recovery of AI applications and databases across cloud and on-premises environments.
What we are looking for
6- Define architecture and technical roadmap for scalable enterprise AI runtime platform
- Design Kubernetes-based systems for AI applications, inference services, and databases
- Build control-plane services, APIs, operators, and automation
- Develop runtime capabilities for GPU scheduling, autoscaling, and load balancing
- Improve performance, availability, and developer experience of AI inference services
- Lead technical initiatives, mentor engineers, and establish platform standards
Skills mentioned
5PythonC++JavaKubernetesCI/CD
