Open role
Service Reliability Engineer
NVIDIA
US, TX, RemotePosted Aug 11, 2026 · 8h ago$168k – $334k
Full-time$168k – $334kSeniorRemoteConsumer Tech
About this role
NVIDIA is seeking a Site Reliability Engineer to ensure high availability for their on-prem and cloud products and services. You will operate within a 24/7 support model, monitor and manage GPU and Kubernetes environments, and develop automated solutions to prevent incidents. This role offers the opportunity to work with cutting-edge technology in AI and accelerated computing.
What we are looking for
6- Operate within a 24/7 follow-the-sun support model
- Monitor and manage GPU and Kubernetes environments for high availability
- Develop predictive automated support routines and auto-healing solutions
- Perform systems administration, network administration, and security monitoring
- 8+ years of experience coordinating large-scale production systems
- Expert-level Linux system administration and automation using Python/Ansible
Skills mentioned
6PythonAWSAzureGCPKubernetesLinux
