Open role
Senior Site Reliability Engineer
NVIDIA
2 LocationsPosted Aug 26, 2026 · 1h ago
Full-timeSeniorHybridConsumer Tech
About this role
NVIDIA is seeking a Principal Staff SRE to lead initiatives in improving core infrastructure services across on-premises and cloud environments. This role involves designing, scaling, and deploying critical infrastructure services, optimizing performance and reliability, and developing tools for data analysis and monitoring.
What we are looking for
6- Lead initiatives to transform IT Compute Core architecture
- Design, scale, and deploy core infrastructure services (DNS, NTP/PTP, DHCP, LDAP)
- Implement service-efficiency metrics and drive improvements
- Apply technologies like eBPF and XDP for observability and DDoS mitigation
- Develop and maintain tools for data analysis, reporting, and monitoring
- 12+ years of experience in compute platform engineering with automation and technical leadership
Skills mentioned
3PythonLinuxTerraform
