At MyFitnessPal, we believe good health starts with what you eat. We provide tools, resources and support to enable users to reach their health goals. We are looking for a Software Engineer III - Site Reliability to join the MyFitnessPal PEAS team. As a member of the PEAS team, you'll have the opportunity to positively impact MyFitnessPal users through your expertise in the reliability and delivery systems that keep the MyFitnessPal ecosystem fast, available, and safe to change. In addition to technical expertise, you'll find that your teammates value collaboration, mentorship, and inclusive environments.
🏢 About MyFitnessPal
The Productivity Engineering: Automation & Self-Service (PEAS) team is responsible for the automation, CI/CD, and self-service platforms that product teams rely on to build, test, and ship features quickly and safely. The PEAS team is part of the Technology Operations (TechOps) organization, which includes IT, Infrastructure, Security, Reliability, and DevOps/DevEx disciplines. TechOps seeks to enable MyFitnessPal to \"ship with confidence\" by delivering a secure, reliable, and low-friction runway to build and operate software at scale.
🎯 The Role
As a Software Engineer III on the PEAS team, you will own reliability across our production services and maintain security in the software delivery pipeline. You'll define how we measure and defend reliability, lead us through incidents, and make our CI/CD pipelines both faster and harder to compromise. You will play a key role in delivering an outstanding experience to MyFitnessPal users.
✅ Key Responsibilities
Own and evolve our SLI/SLO and error-budget frameworks, and use them to influence prioritization and product decisions
Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches
Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue
Design and operate resilient, scalable infrastructure using Infrastructure as Code (Terraform)
Manage production Kubernetes and container workloads, including capacity planning and cloud-cost optimization
Own CI/CD pipelines and safe deployment strategies (canary, progressive rollout, fast rollback)
Own the security controls that live inside the delivery pipeline — integrating and tuning SAST, DAST, and SCA scanning (for example, in GitHub Actions) so issues surface while code is still in review
Implement and maintain policy-as-code (for example, OPA/Rego, Kyverno, or Conftest) to block unsafe infrastructure and Kubernetes changes at admission time
Drive vulnerability triage and remediation SLAs for pipeline- and infrastructure-level findings, prioritizing by real risk
Partner with our Security Engineer and the broader Security & Reliability disciplines — you own security in the pipeline and collaborate on the rest, rather than duplicating that function
Participate in and improve the on-call rotation; build the runbooks and automation that make on-call sustainable
Coach team members and engineers across the org on reliability patterns and operational best practices
📌 Required Qualifications
5+ years in site reliability, platform, or infrastructure engineering, with clear senior-level ownership of production systems
Strong programming skills for automation and tooling (Go, Python, Typescript or similar) — you have experience building software or custom tooling, not just scripts
Deep, hands-on experience with a major cloud platform (AWS is a plus), Kubernetes, and Infrastructure as Code (Terraform is a plus)
Proven track record leading incident response and building SLO-driven reliability practices.
Working fluency with observability tooling (Datadog is a plus)
Practical experience integrating security into CI/CD pipelines — SAST/DAST/SCA tooling, dependency scanning, or policy-as-code
The judgment and communication skills to raise a security or reliability finding with a senior engineer and land it as a shared problem to solve, not a fight to win
⭐ Desirable Experience
Experience with policy-as-code frameworks (especially Kyverno, but tools like OPA/Rego or Conftest are also relevant) enforced at admission time is a plus
Exposure to regulated or compliance-driven environments (SOC 2, PCI DSS, HIPAA) is a plus
Chaos engineering or game-day experience is a plus
Experience supporting B2C/mobile backend environments with high traffic, rapid iteration, and strong reliability needs is a plus
🎁 Benefits
The reasonably estimated salary for this role at MyFitnessPal ranges from $120,000 - $165,000. Actual compensation is based on factors such as the candidate’s skills, qualifications, and experience. In addition, MyFitnessPal offers a wide range of comprehensive and inclusive employee benefits for this role including healthcare, parental planning, mental health benefits, annual performance bonus, a 401(k) plan and match, responsible time off, monthly wellness and technology allowances, and others.
Please let Myfitnesspal know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.