Open role
Senior Site Reliability Engineer - Storage
NVIDIA
US, CA, Santa ClaraPosted Sep 1, 2026 · 3h ago$168k – $334k
Full-time$168k – $334kSeniorOn-site
About this role
NVIDIA is seeking a Senior Site Reliability Engineer to design, implement, and optimize on-prem High-Performance Computing (HPC) storage solutions, integrating cloud technologies. You will build automation tools, ensure efficient operations, and collaborate with engineering teams to meet infrastructure needs for AI and HPC projects.
What we are looking for
6- Design and implement on-prem HPC storage infrastructure with cloud integration
- Develop scalable storage solutions for data-intensive applications
- Automate infrastructure deployment, management, and monitoring
- Evaluate and document distributed file systems and best practices
- Collaborate with engineering teams to gather infrastructure requirements
- Experience with enterprise NAS, S3 storage, and parallel/distributed file systems
Skills mentioned
7PythonGoAWSAzureGCPDockerKubernetes
