Open role
Staff Site Reliability Engineer - AI Platform Runtime
NVIDIA
US, CA, Santa ClaraPosted Sep 9, 2026 · 6h ago$168k – $334k
Full-time$168k – $334kStaffHybridAI / ML
About this role
NVIDIA is seeking a Staff Site Reliability Engineer to lead technical strategy and roadmap for large-scale SRE initiatives, focusing on reliability, scalability, and developer productivity for AI-driven enterprise products. You will design and build resilient distributed systems, architect AI Agents, and drive automation and observability improvements.
What we are looking for
6- Lead technical strategy and roadmap for SRE initiatives
- Design and build resilient distributed systems for AI products
- Architect and develop AI Agents and Skills for platform operations
- Drive automation and observability improvements
- Collaborate across Cloud, Platform, Security, and AI/ML teams
- Mentor and influence engineers across teams
Skills mentioned
8PythonJavaScriptTypeScriptAWSAzureGCPKubernetesTerraform
