Open role
Sr. Site Reliability Engineer
Clearwateranalytics
Office - BoisePosted Sep 24, 2026 · 5h ago
SeniorOn-siteSaaS
About this role
Clearwater Analytics is seeking a Senior Site Reliability Engineer to ensure the performance, scalability, and reliability of our cloud-native systems. You will drive automation, implement robust monitoring, and lead incident management using Kubernetes and observability tools. This role is key to maintaining operational excellence and improving system reliability through innovative solutions, including AI-assisted troubleshooting.
What we are looking for
6- Design, build, and maintain highly available, scalable, and reliable production systems
- Automate infrastructure provisioning and operations using Terraform (IaC)
- Operate and manage cloud-native platforms, including Amazon EKS
- Implement and maintain monitoring, logging, and alerting using Prometheus, Grafana, Dynatrace, and OpenSearch
- Lead incident management, production troubleshooting, and root cause analysis (RCA)
- Drive AI-assisted investigations as a core part of incident response
Skills mentioned
9PythonJavaSQLAWSDockerKubernetesLinuxTerraformCI/CD
