Open role
Senior Site Reliability Engineer - Storage
NVIDIA
US, CA, Santa ClaraPosted Aug 10, 2026 · 15h ago$168k – $322k
Full-time$168k – $322kSeniorOn-siteEnterprise Software
About this role
NVIDIA is seeking a Senior Site Reliability Engineer to design, implement, and optimize on-prem High-Performance Computing (HPC) storage solutions, integrating cloud technologies. You will build automation tools, ensure efficient operations, and collaborate with engineering teams to meet evolving infrastructure needs for groundbreaking AI and HPC projects.
What we are looking for
6- Design and implement on-prem HPC storage infrastructure with cloud integration
- Develop scalable and efficient storage solutions for data-intensive applications
- Automate deployment, management, and monitoring of large-scale infrastructure
- Evaluate and document procedures for distributed file systems
- Collaborate with engineering teams to gather infrastructure requirements
- Experience with enterprise NAS, S3 storage, and parallel/distributed file systems
Skills mentioned
7PythonGoAWSAzureGCPDockerKubernetes
