Open role
Senior Staff Site Reliability Engineer
NVIDIA
India, BengaluruPosted Aug 12, 2026 · 4h ago
Full-timeStaffOn-siteAI / ML
About this role
NVIDIA is seeking a Senior Staff Site Reliability Engineer to lead technical strategy for large-scale SRE initiatives, focusing on reliability, scalability, and developer efficiency. You will design and build resilient distributed systems for AI-powered enterprise products, drive automation and observability improvements, and develop LLM-aware monitoring and incident response pipelines.
What we are looking for
6- Lead technical strategy for SRE initiatives
- Design and build resilient distributed systems for AI products
- Drive automation and observability improvements
- Build LLM-aware monitoring and autonomous incident response
- Collaborate with Cloud, Platform, Security, and AI/ML groups
- Analyze and run complex systems including Kubernetes and AI/ML infrastructure
Skills mentioned
8PythonJavaScriptTypeScriptAWSAzureGCPKubernetesTerraform
