Open role
Principal Site Reliability Engineer, Machine Learning
Cambridge Mobile Telematics
Cambridge, MAPosted Jul 27, 2026 · 6h ago$142k – $178k
Full-time$142k – $178kPrincipalHybridAI / ML
About this role
Cambridge Mobile Telematics is seeking a Principal Site Reliability Engineer to ensure the operational health and scalability of their machine learning platforms on AWS. This role involves managing EKS and Databricks workloads, implementing infrastructure as code, and leading incident response to enhance road safety through AI.
What we are looking for
6- Own SLOs, error budgets, and operational health of AWS EKS and Databricks workloads
- Maintain observability and alerting using CloudWatch and Datadog
- Operate and tune EKS Ray workloads at scale, including autoscaling and GPU scheduling
- Manage Databricks on AWS, including workspace administration and job orchestration
- Implement infrastructure-as-code using Terraform and CI/CD pipelines
- Lead incident response and drive systemic remediation for platform outages
Skills mentioned
6PythonAWSDockerKubernetesTerraformMachine Learning
