Hims & Hers is the leading health and wellness platform, on a mission to help the world feel great through the power of better health. We are redefining healthcare by putting the customer first and delivering access to care that is affordable, accessible, and personal, from diagnosis to treatment to delivery. No two people are the same, so we provide access to personalized care designed for results. By normalizing health & wellness challenges and innovating on their solutions, we’re making better health outcomes easier to achieve. Hims & Hers is a public company, traded on the NYSE under the ticker symbol “HIMS.”
🏢 About Hims & Hers
Hims & Hers is the leading health and wellness platform, on a mission to help the world feel great through the power of better health. We are redefining healthcare by putting the customer first and delivering access to care that is affordable, accessible, and personal, from diagnosis to treatment to delivery. No two people are the same, so we provide access to personalized care designed for results. By normalizing health & wellness challenges and innovating on their solutions, we’re making better health outcomes easier to achieve. Hims & Hers is a public company, traded on the NYSE under the ticker symbol “HIMS.” To learn more about the brand and offerings, you can visit hims.com/about and hims.com/how-it-works . For information on the company’s outstanding benefits, culture, and its talent-first flexible/remote work approach, see below and visit www.hims.com/careers-professionals.
🎯 The Role
We're looking for a Senior Reliability Engineer to make Hims & Hers systems measurably more reliable: building the observability, tooling, and automation that catch problems before people do. Recent work at this level includes preparing our stack for 32x baseline traffic during high-stakes seasonal surges with zero customer-facing issues and improved P95 latency, turning databases into a paved platform with consistent observability and guardrails, and building AI agents that pull FireHydrant, Datadog, Jira, and Confluence into a single incident picture. We use AI first: every engineer gets a Claude Enterprise license, and we expect you to use it as a core part of how you investigate, build, and write, not as an afterthought.
✅ Key Responsibilities
Own reliability for Tier 1 customer journeys: Define and instrument SLOs, golden signals, and business-level monitors for the journeys that matter most (checkout, telehealth visits, prescription fulfillment), so that degradation is detected by systems rather than by customers or support tickets. Move our \"detected by monitors vs. detected by humans\" ratio in the right direction and be able to show it.
Engineer for peak load and failure: Lead capacity and resilience work for high-stakes events and steady-state growth: load-test design, second-by-second analysis of prior events, database and backend bottleneck investigation, and hardening across caching, GraphQL, VPC capacity, and vendor rate limits. Validate before game day, not during it.
Build incident response as software: Mature FireHydrant, Datadog, and Jira into one connected pipeline: automated incident and RCA ticket creation, SLO burn-rate and composite alerting with deduplication, Tier 1 alert routing, and enforced post-mortem action tracking. Author and maintain the runbooks and severity standards that make any responder effective on any service.
Automate operational excellence with AI: Build and operate agents and tooling that reduce manual OE work: OER report generation, RCA drafting, stale action-item detection, monitor and runbook gap detection, and OpenClaw agents wired to Datadog and FireHydrant for first-pass incident triage. Ship these as reusable capabilities, not personal scripts.
Be the deep-debugging expert, and make teams better at it: Teams own debugging their own services, but you are the person they pull in when a problem crosses boundaries: silent service-to-service failures, anomalous traffic, or regressions that span frontend, API, mesh, and database layers. Bring that depth to the hardest cases yourself, then turn what you learn into runbooks, tooling, and pairing so teams can catch and resolve the next one on their own. Partner with Security and product teams on anomaly detection and response, weighing engineering cost and user impact alongside the benefit of any control.
Maintain the tooling that informs Operational Excellence reviews: Produce the tooling that powers the metrics and narrative that go into bi-weekly VP-level OE reviews and the monthly cross-engineering OER, and drive the resulting action items to closure.
Raise the bar for others: Document what you build, onboard teammates to roll it out, and coach engineers across squads on SLOs, blameless post-mortems, and on-call practice.
📌 Required Qualifications
5+ years as a Software, SRE, Platform, or Infrastructure Engineer, with a track record of owning reliability outcomes for production systems that customers depend on.
Strong software engineering fundamentals. You solve reliability problems by writing code and building tooling, and you're comfortable reading application code across the stack to find the real cause.
Hands-on depth in observability and SLO engineering: golden signals, burn-rate alerting, journey-level monitors, and turning noisy alert streams into actionable pages (Datadog preferred; Prometheus/Grafana and OpenTelemetry welcome).
Production experience with AWS, Kubernetes/EKS, Terraform, and PostgreSQL (RDS/Aurora).
Experience running or maturing incident management end to end: on-call design, escalation policies, incident command, blameless post-mortems, and action-item follow-through in a tool like FireHydrant or PagerDuty.
⭐ Desirable Experience
Daily, practical use of AI coding and analysis to
🎁 Benefits
Outlined above is a reasonable estimate of H&H’s compensation range for this role for US-based candidates. If you're based outside of the US, your recruiter will be able to provide you with an estimated salary range for your location. The actual amount will take into account a range of factors that are considered in making compensation decisions, including but not limited to skill sets, experience and training, licensure and certifications, and location. H&H also offers a comprehensive Total Rewards package that may include an equity grant.
🛂 Visa & Eligibility
Consult with your Recruiter during any potential screening to determine a more targeted range based on location and job-related factors.
Please let Hims & Hers know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.