Software Development👥 1001 employees📍 San Francisco, California, USEst. 2004
Delinea is a pioneer in securing identities through centralized authorization, making organizations more secure by seamlessly governing their interactions across the modern enterprise.
Delinea al…
📋 Job Overview
Our growing technology company is seeking an experienced Site Reliability Engineer with deep Azure expertise to help maintain the availability, performance, and reliability of our critical SaaS applications. In this role, you will own and drive automation, monitoring, incident response, and infrastructure improvements across our multi-cloud, multi-region environment, working closely with senior engineering and cross-functional teams.
🏢 About Delinea
Delinea is a pioneer in securing human and machine identities through intelligent, centralized authorization, empowering organizations to seamlessly govern their interactions across the modern enterprise. Leveraging AI-powered intelligence, Delinea’s leading cloud-native Identity Security Platform applies context throughout the entire identity lifecycle – across cloud and traditional infrastructure, data, SaaS applications, and AI. It is the only platform that enables you to discover all identities – including workforce, IT administrator, developers, and machines – assign appropriate access levels, detect irregularities, and respond to threats in real-time. With deployment in weeks, not months, 90% fewer resources to manage than the nearest competitor, and a 99.995% uptime, Delinea delivers robust security and operational efficiency without compromise. Learn more about Delinea on Delinea.com, LinkedIn, X, and YouTube.
🎯 The Role
Join our passionate, global team at Delinea and help us make the world a safer and more secure place. Our success is driven by world-class product leadership, outstanding engineers, and strategic investment from TPG. We value diversity, innovation, and a culture of respect and fairness. If you're ready to push boundaries and challenge the status quo in security, we want to hear from you.
✅ Key Responsibilities
Own the availability and performance of production SaaS applications running on Azure (AKS, App Service, Redis, SQL, Service Bus etc), across multiple geographic regions.
Lead troubleshooting and resolution of cloud infrastructure and application issues, including AKS pod/node failures, deployment rollbacks, ingress and networking issues, and resource/autoscaling problems.
Participate in an on-call rotation (including weekends) and drive incident response from detection through resolution, with a primary focus on customer experience and minimizing impact.
Drive improvements to disaster recovery, failover, and incident management processes across multi-region deployments.
Build and maintain automation scripts and monitoring tools to reduce manual toil and streamline operational tasks.
Author post-incident reviews (RCAs), identify root causes, and drive preventive action items to closure.
Partner with senior engineers and cross-functional teams to implement and improve reliability, observability, and performance best practices.
Contribute to continuous improvement initiatives across infrastructure, tooling, and process.
Communicate clearly with customer-facing stakeholders when incidents require external status updates or written incident summaries.
📌 Required Qualifications
5+ years of relevant experience in Site Reliability Engineering, DevOps, or Cloud Administration, with demonstrated ownership of production systems.
Hands-on experience administering Azure environments, including AKS (Kubernetes), core Azure services, cloud networking, and cloud security fundamentals.
Solid understanding of monitoring, logging, and alerting practices (e.g., Datadog, Azure Monitor, ELK stack), including hands-on troubleshooting with log analysis and stack traces using Datadog APM.
Familiarity with networking fundamentals: firewalls, load balancers, VPNs, DNS, and routing.
Experience with automation and scripting (PowerShell, Python, or similar).
Practical understanding of backup, redundancy, and disaster recovery strategies in cloud environments, including geo-redundant / multi-region deployments.
Strong ownership mindset across the full incident lifecycle, from detection through post-mortem, with a customer-first approach.
Comfort communicating clearly and professionally in written form, including incident status updates and post-incident summaries.
⭐ Desirable Experience
Experience with AWS Cloud Platform.
Experience with CI/CD tools such as Azure DevOps.
Experience with infrastructure-as-code tools such as Terraform or ARM templates.
Prior experience operating SaaS products with regional tenant architectures (e.g., multiple geo-specific production environments).
🎁 Benefits
We offer competitive salaries, a meaningful bonus program, and excellent benefits, including healthcare insurance, as well as pension/retirement matching, comprehensive life insurance, an employee assistance program, time off plans, and paid company holidays.
Please let Delinea know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.