Open role
Site Reliability Engineer
nscaleoperationsukltd
Houston; New York; San Francisco; SeattlePosted Sep 8, 2026 · 19h ago$130k – $200k
Full-time$130k – $200kMid LevelHybridSaaS
About this role
Nscale is seeking a Site Reliability Engineer to own and improve the automation and tooling for their GPU cloud infrastructure. You will be responsible for the reliability of production services, incident response, and driving system improvements through code. This role offers flexibility and the opportunity to make a significant impact on AI infrastructure.
What we are looking for
6- Build and own automation and tooling for platform reliability
- Define and maintain SLOs, SLIs, and dashboards
- Lead incident response, troubleshooting, and root cause analysis
- Improve availability, scalability, and efficiency through code
- Partner with Engineering, Networking, and Infrastructure teams
- Experience with Python, Kubernetes, Linux, and distributed systems
Skills mentioned
2PythonKubernetes
