Open role
Staff Site Reliability Engineer
Wonder
Tel AvivPosted Sep 29, 2026 · 4h ago
StaffFood
About this role
As a Staff Site Reliability Engineer at Wonder, you will architect and implement resilient, self-healing systems to enhance the student dining experience across the US. You'll manage AWS infrastructure, drive observability improvements, design scaling strategies, and shape incident management processes, ensuring our platform scales to meet growing demand.
What we are looking for
6- Architect resilient, self-healing systems and co-own critical production services
- Own multi-region resilience, including failover design and data-layer replication
- Manage AWS infrastructure as code using Terraform/Terraspace
- Operate and upgrade the Kubernetes platform (EKS, Helm, autoscaling)
- Own the end-to-end observability platform (logging, metrics, tracing, alerting)
- Drive reliability improvements using SLOs and telemetry data
Skills mentioned
8PythonAWSKubernetesLinuxMongoDBRedisTerraformCI/CD
