Rocket Money’s mission is to empower people to live their best financial lives. Rocket Money offers members a unique understanding of their finances and a suite of valuable services that save them time and money – ultimately giving them a leg up on their financial journey.
🏢 About Rocket Money
We're looking to expand our Cloud Infrastructure team with a Senior Infrastructure Engineer, SRE to lead the reliability and operational evolution of our platform. We run hundreds of services in production, which enable us to process billions of transactions, consume multiple terabytes of data, and produce hundreds of millions of logs per day, and our reliability practice needs to evolve to match our growing scale.
🎯 The Role
You'll join the Cloud Infrastructure team and partner with engineering and internal support teams to drive this work. We support millions of people to improve their financial lives, and this role ensures we can continue to do so reliably and at scale.
✅ Key Responsibilities
Building and improving the reliability and resiliency of our systems and services
Establishing SLIs, SLOs, and error budgets for our most critical services and user journeys, and reviewing them regularly with the teams that own them
Owning and evolving our disaster recovery strategy: recovery objectives, failover and restore paths, and regular exercises that prove they work
Partnering with product engineering teams so they can own and operate their own services, with metrics that reflect real user experience
Evolving our observability platform and standards across metrics, tracing, and logs: including instrumentation paved roads, alert quality, and observability cost
Strengthening our incident practice: tuning paging thresholds, keeping runbooks current, and following through on postmortem action items
Contributing to day-to-day Cloud Infrastructure work alongside your reliability specialty — infrastructure build-outs, platform backlog, and a shared on-call rotation (1 week out of every 6 weeks)
📌 Required Qualifications
You have 5+ years of hands-on cloud or infrastructure engineering experience, with substantial time spent on reliability and production operations at scale
You have defined SLIs and SLOs for real production services, and can talk about what changed as a result. What got fixed, what got deprioritized, and what you got wrong the first time
You have hands-on experience with an observability platform in production; Datadog strongly preferred
You're comfortable writing code (Python, Go, TypeScript, or similar) for internal tooling, production debugging, and automation
You write production Terraform and are comfortable in AWS, and when production breaks you can find the problem and fix it
You have built or operated a disaster recovery plan: you set the recovery goals, wrote the failover and restore steps, and ran the drills that proved it works
You have been on-call for services you helped build, and you have opinions about what makes an alert worth waking someone for
You prefer giving teams paved roads and good defaults over mandates, so they can own their own instrumentation
⭐ Desirable Experience
You have led a reliability or observability modernization project where you defined the vision, approach, and delivered the implementation
You have built internal tooling, libraries, or instrumentation standards that made it easier for other teams to operate their services well
You have run game days, chaos experiments, or DR exercises, and fixed the problems they uncovered
You have cut observability spend while keeping the coverage you needed
🎁 Benefits
Health, Dental & Vision Plans
Competitive Pay
401k Matching
Unlimited PTO
Lunch daily (in-office only)
Snacks & Coffee (in-office only)
Commuter benefits (in-office only)
🛂 Visa & Eligibility
Los Angeles County and San Francisco Candidates only: qualified applicants with arrest or conviction records will be considered for employment per the Fair Chance Ordinance and the Fair Chance Initiative for Hiring.
Please let Rocket Money know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.