Level is a learning technology company dedicated to helping students build real academic and life skills with confidence and joy. We combine proven curriculum principles with world class interactive design to make meaningful practice something students want to come back to, not something they struggle through. We support what teachers, schools, and parents are already doing by increasing student engagement with high quality, standards aligned practice that reinforces classroom learning.
🏢 About Level
As an Senior Infrastructure Engineer on the Platform team, you will architect, build, and operate the cloud infrastructure and developer platform that every Level product runs on. You will own critical infrastructure end-to-end — the Kubernetes platform, infrastructure-as-code, CI/CD and GitOps delivery, networking (ingress and egress), DNS, observability, and cloud security posture — and provide the reliable, self-service foundations the rest of engineering builds on. You will work on a small, senior-leaning team where infrastructure decisions have direct, visible impact on reliability, performance, cost, and developer velocity.
🎯 The Role
You will be responsible for designing, building, and operating secure, highly available AWS infrastructure using Terraform/OpenTofu with a GitOps workflow (Atlantis). You will own capacity planning, DR, and cost optimization for the systems you run. You will operate and evolve EKS: autoscaling (Karpenter), upgrades, core add-ons, and Helm-based delivery (ArgoCD). You will build and maintain GitHub Actions pipelines that let platform and product teams ship fast and safely, with self-service tooling where it makes sense.
✅ Key Responsibilities
Design, build, and operate secure, highly available AWS infrastructure using Terraform/OpenTofu with a GitOps workflow (Atlantis).
Own capacity planning, DR, and cost optimization for the systems you run.
Operate and evolve EKS: autoscaling (Karpenter), upgrades, core add-ons, and Helm-based delivery (ArgoCD).
Build and maintain GitHub Actions pipelines that let platform and product teams ship fast and safely, with self-service tooling where it makes sense.
Own ingress/egress (Traefik), service mesh and mTLS (Linkerd/Envoy), load balancing, edge TLS, and DNS (Route 53, Terraform-managed).
Build observability with OpenTelemetry and SigNoz; use telemetry to drive reliability, performance, and cost decisions.
Serve as an escalation point for complex incidents, leading troubleshooting and post-mortems.
Apply cloud security best practices across identity, secrets, and network boundaries, with particular care for student data and K-12 privacy.