Open role
Principal Site Reliability Engineer
NVIDIA
US, CA, Santa ClaraPosted Sep 9, 2026 · 6h ago
Full-timePrincipalOn-siteConsumer Tech
About this role
NVIDIA is seeking a Principal Site Reliability Engineer to lead the technical direction and reliability initiatives for their AI Platform Runtime. This role involves architecting highly available distributed systems, developing AI-driven automation for platform operations, and establishing reliability standards across the organization. You will influence technical strategy, mentor engineers, and drive improvements in availability, scalability, and performance.
What we are looking for
6- Define and drive long-term technical vision for reliability across NVIDIA's AI Platform Runtime
- Architect highly available, resilient, secure, and scalable distributed platforms
- Lead design and development of AI agents and intelligent automation for platform operations
- Establish platform-wide reliability standards including SLOs and capacity models
- Identify systemic risks and lead cross-functional programs to improve system reliability
- Provide technical leadership during critical incidents and drive improvements
Skills mentioned
11PythonJavaJavaScriptTypeScriptAWSAzureGCPKubernetesLinuxTerraformMachine Learning
