Hospitals and Health Care👥 51 employees📍 New York, New York, USEst. 2021
Certify is a first-of-its-kind provider intelligence platform, powered by API integrations and hundreds of verified data points. We unlock insights and power performance for clinicians, teams and org…
📋 Job Overview
We’re looking for a Senior Site Reliability Engineer who takes ownership seriously — someone who designs for reliability, ships the automation, and stands behind it in production. You’ll work across cloud-native infrastructure on systems that process millions of provider records.
🏢 About CertifyOS
CertifyOS is building the data infrastructure that powers modern healthcare. Today, healthcare organizations rely on fragmented and outdated provider data. This creates unnecessary administrative work, regulatory risk, and higher costs across the system. We’re solving that problem.
Our API-first platform automates provider licensing, enrollment, credentialing, and network monitoring by connecting directly to hundreds of primary data sources. We help healthcare organizations maintain accurate, compliant, and reliable provider networks at scale.
Our vision is simple: One API. One provider ID. Frictionless provider data.
We’re backed by leading investors and built by a team with deep experience in provider data systems. At CertifyOS, we value authenticity, accountability, collaboration, results, and openness to feedback. We’re building a high-ownership team focused on solving real infrastructure problems that impact millions of patients.
🎯 The Role
This is a role with real scope: you’ll own the operational lifecycle end-to-end and influence platform architecture, reliability standards, and deployment workflows across systems that matter.
✅ Key Responsibilities
Reliability and observability at scale: maintain uptime, reduce alert fatigue, and build actionable observability across GKE and Cloud Run
Scaling infrastructure efficiently: improve autoscaling behavior, resource utilization, and workload efficiency across cloud-native distributed systems
Incident response and operational maturity: own incident response processes, root cause analysis, escalation workflows, and runbooks
Infrastructure automation and developer velocity: build and maintain Infrastructure as Code, CI/CD pipelines, and operational tooling
Reliability engineering for data platforms: instrument data freshness and infrastructure health
📌 Required Qualifications
5+ years in SRE, DevOps, Platform Engineering, or Infrastructure Engineering — operating production systems at scale
Track record of improving reliability end-to-end: debugged hard production problems and made them not happen again
Strong Linux systems administration, incident response, and root cause analysis skills
Deep hands-on experience with GCP — GKE, Cloud Run, and containerized workloads at scale
Experience building and maintaining Infrastructure as Code with Terraform and/or Pulumi
Fluency across deployment patterns and the judgment to know when each fits: rolling deployments, blue/green, canary — and the rollback story for each
Experience with autoscaling, resource optimization, and infrastructure efficiency for distributed systems
Experience managing infrastructure security, secrets, and access controls in regulated or security-conscious environments
Strong understanding of Golden Signals monitoring — latency, traffic, errors, saturation — and how to make them actionable
Hands-on experience with observability platforms: Google Cloud Monitoring, Datadog, Grafana, Prometheus, or similar
Experience building and maintaining CI/CD pipelines using GitHub Actions or similar
Scripting or programming fluency in Python, Bash, Go, or similar
Strong written and verbal communication
Experience operating systems handling sensitive data or PII in regulated or compliance-adjacent environments
Please let Certify know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.