Open role
Senior Staff Site Reliability Engineer
NVIDIA
India, BengaluruPosted Sep 23, 2026 · 19h ago
StaffAI / ML
About this role
NVIDIA is seeking a Senior Staff Site Reliability Engineer to build and lead the runtime foundation for our enterprise AI platforms. You will architect and implement scalable Kubernetes-based systems for deploying, operating, and scaling AI applications and databases across cloud and on-premises environments. This role offers a unique opportunity to drive technical strategy, mentor engineers, and shape the future of AI infrastructure.
What we are looking for
6- Define architecture and technical roadmap for enterprise AI runtime platform
- Design Kubernetes-based systems for AI applications and databases
- Build control-plane services, APIs, and automation for workload management
- Develop runtime capabilities for GPU scheduling, autoscaling, and load balancing
- Improve performance and availability of large-scale AI inference services
- Lead technical initiatives and mentor engineers
Skills mentioned
5PythonC++JavaKubernetesCI/CD
