IT Services and IT Consulting👥 51 employees📍 San Francisco, CA, USEst. 2012
Technology challenges can be a headache. We help you tackle your biggest DevOps and IT challenges, delivering superior outcomes you didn’t think were possible. Our teams work with innovative companie…
📋 Job Overview
EverOps is looking for a Lead Kubernetes Platform Engineer with deep Amazon EKS experience at very large scale to lead a modernization discovery, and the engineering program that follows it, for a high-scale consumer mobile platform. You’ll be stepping into a multi-cluster production EKS estate where the largest clusters run tens of thousands of pods at peak and compute demand roughly doubles between overnight lows and daytime highs. Upgrading the estate is slow and largely manual, so the team is perpetually behind the Kubernetes release cycle. Compute is one of the largest lines in the business, and the platform backs real-time, safety-critical features where failing to scale at peak is not an option.
🏢 About EverOps
Enter EverOps – the premier Embedded Service Provider. We partner directly with customer engineering teams to assess and address mission-critical infrastructure, cloud, and delivery challenges.
🎯 The Role
As a Lead Kubernetes Platform Engineer, you will join our U.S.-Based Virtual Operating Center and lead an embedded TechPod as its Pod Leader: a player-coach who sets technical direction, serves as the primary point of contact for the customer’s infrastructure leadership, and stays hands-on in the work. Your immediate priority is leading a two-month EKS modernization discovery. You’ll baseline the production estate, map Kubernetes version and support status cluster by cluster, analyze blast radius and failure domains, and deliver the upgrade automation design, compute and capacity economics model, and prioritized roadmap needed to commit the program.
✅ Key Responsibilities
Estate Discovery: Build a complete baseline of a multi-cluster production EKS estate, covering cluster inventory, topology, workload placement, ownership, per-cluster cost, and Kubernetes version and support status.
Blast Radius & Architecture: Analyze failure domains, isolation boundaries, and dependency concentration in very large clusters, and design a multi-cluster target architecture with clear workload placement and tenancy models.
Upgrade Automation: Design and implement an upgrade approach (in-place vs. blue/green clusters, Karpenter drift-based node rotation, add-on and API deprecation management) that turns a months-long manual cycle into a repeatable, largely automated process.
Compute & Capacity Strategy: Lead instance sizing and workload-fit analysis across instance families, generations, and node sizes, accounting for DaemonSet overhead, bin-packing, network limits, and headroom for rapid scale-up.
Graviton Migration: Plan and drive a phased ARM64 migration across Java, Go, PHP, and Python workloads, including multi-architecture builds, native dependency remediation, and per-service rightsizing against latency SLOs.
Karpenter Engineering: Configure NodePools, weights, and node overlays so scheduling favors price-performance rather than hourly price, and design safe Spot patterns that protect interruption-sensitive and stream-processing workloads.
Cost Engineering: Model compute, support, and commitment economics (Savings Plans, On-Demand, Spot suitability by workload tier, extended support exposure) and quantify savings ranges by lever.
Platform Posture: Assess and improve ingress (including migration off Ingress NGINX toward Gateway API), service mesh, CNI, GitOps, and infrastructure-as-code posture across the estate.
Operability & Toil Reduction: Quantify maintenance and toil burden using the customer’s own engineering data, set reduction targets, and build the automation that gives time back to platform teams.
AI-Native Operations: Define the scope and success criteria for an AI-assisted DevOps agent workstream, and assess platform readiness for AI-accelerated development demand.
Technical Leadership: Lead the TechPod as a player-coach, run the working cadence with the customer’s infrastructure owner, and partner with AWS specialists on capacity planning and architecture decisions.
Documentation & Readouts: Produce the estate baseline, architecture designs, upgrade plans, roadmap, and executive readout, and present findings and recommendations to engineering leadership.
📌 Required Qualifications
Experience: 8+ years in DevOps, SRE, Platform, or Infrastructure Engineering, including 4+ years operating production Kubernetes and prior experience in a technical lead, staff, or principal-level role.
EKS at Scale: Deep production experience with Amazon EKS at large scale (multiple production clusters, thousands of nodes, or
Please let EverOps know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.