Horizon3 is a fast-growing, remote cybersecurity company dedicated to the mission of enabling organizations to proactively find, fix and verify exploitable attack vectors before criminals exploit them. Our flagship product, the NodeZeroTM platform, delivers production-safe autonomous pentests and other key assessment operations that scale across the largest internal, external, cloud, and hybrid cloud environments. NodeZero has been adopted by organizations of all sizes, from small educational institutions to government agencies and Global 100 enterprises. It is used by IT Ops/SecOps teams, consulting pentesters, and MSSPs and MSPs.
🏢 About Horizon3
We are a fusion of former U.S. Special Operations cyber operators, startup engineers & operators, and formerly frustrated cybersecurity practitioners. We're committed to helping solve our common security problems: ineffective security tools and false positives, resulting in alert fatigue, blind spots, \"checkbox” security culture, cybersecurity skills shortage, and the long lead time and expense of hiring outside consultants. Collectively, we are a team of learn it alls, committed to a culture of respect, collaboration, ownership, and results.
🎯 The Role
We are seeking a hands-on Staff Site Reliability Engineer to own and evolve the reliability strategy, operating model, and engineering-wide standards supporting our platform. This is a foundational role for an experienced engineer who will set technical direction across teams, lead the highest-impact reliability initiatives, and establish the practices and systems that enable engineering to operate production services safely.
✅ Key Responsibilities
Own and evolve the engineering-wide SRE strategy, operating model, and reliability standards, aligning them to customer impact, business priorities, and risk.
Lead cross-functional alignment across Infrastructure, product, service, security, and business stakeholders to improve reliability, observability, incident response, and operational readiness across multiple teams.
Establish an organization-wide approach to service ownership, meaningful SLIs and SLOs, and error budgets for critical customer paths and services.
Define and drive adoption of observability standards across pipelines and platform components that report on service health, performance and operational risk.
Set the standard for dashboards, actionable alerting, runbooks, and escalation paths.
Drive end to end complex cross-functional reliability initiatives.
Set and raise the engineering wide bar for incident management, incident command, on-call health, post-incident learning, and recovery readiness.
Shape the technical direction, operationable model, and growth path of the SRE function.
Participate in a 24/7 on-call rotation and help design an on-call model that is sustainable, appropriately staffed, and continuously improved.
📌 Required Qualifications
Experience designing, operating, and troubleshooting large scale distributed systems in production environments.
Deep knowledge of reliability engineering, observability, incident management, and production operations, with demonstrated ability to turn that knowledge into standards and practices adopted by others
Experience in establishing SLIs, SLOs, actionable alerts, observability, and service ownership.
Backend experience building backend systems and automation that reduce optional toil, strengthen safeguards, and operational efficiency.
Experience in leading high severity incidents and improving incident response programs.
Excellent written and verbal communication skills including technical designs, runbooks, postmortems, and operational documentation.
Required Tech Stack Experience: Python and Terraform (Infrastructure as Code), or equivalent automation and infrastructure-as-code tools.
Experience with Observability tools such as Datadog, New Relic, Grafana, or equivalent platforms.
Experience operating production services in AWS and Kubernetes
Experience with CI/CD pipelines such as Gitlab CI, ArgoCD, or GitOps workflows..
🎁 Benefits
Inclusive Team: We value diversity and promote an inclusive culture where everyone can thrive.
Growth Opportunities: Be part of a dynamic and growing team with numerous career development opportunities.
Innovative Culture: Work in a collaborative environment that encourages creativity and out-of-the-box thinking.
Hybrid & Remote Work: We embrace a mix of remote and hybrid work models depending on role and location, including our Chicago office, where some roles require regular in-office presence.
Competitive Compensation: We offer competitive salary, equity and benefits. Our benefits include health, vision & dental insurance for you and your family, a flexible vacation policy, and generous parental leave.
🛂 Visa & Eligibility
We are a fully remote company, and this job may require up to 10% of travel to be successful. Travel primarily consists of team off-sites and in-person project kick-offs.
Please let Horizon3.ai know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.