The challenge
As a Site Reliability Engineer, you will play a critical role in ensuring the stability, availability, resilience, and performance of our cloud platform in production.
Your mission will cover two complementary areas. First, you will work proactively to prevent incidents by identifying operational risks, anticipating capacity constraints, improving production changes, resolving recurring problems, and driving reliability improvements, migrations, and corrective initiatives. Second, when incidents occur, you will help restore service quickly by coordinating technical diagnosis, controlling the situation, communicating impact, and leading recovery efforts.
This is a senior, hands-on role for someone with a strong operational mindset who is comfortable working in complex production environments, investigating technical issues, making decisions under pressure, and taking ownership of platform reliability.
You will be part of the Site Reliability Engineering team, working in a practical environment focused on keeping the platform stable while continuously reducing operational risk and manual effort.
A key part of the role will be to identify what could affect the platform before it becomes an incident. You will assess risks, make them visible, define mitigation or correction plans, assign ownership, and follow actions through to completion. You will also contribute to capacity planning, production scaling, infrastructure improvements, migrations, observability, automation, and operational readiness.
When incidents happen, you will act as a technical reference and coordination point. You will help establish control, assess impact, accelerate diagnosis, involve the right teams, communicate clearly, and restore service as quickly and safely as possible.
You will work closely with engineering, systems, networking, storage, database, security, and product teams. Your operational perspective will help ensure that reliability, scalability, recoverability, and operability are considered in platform changes and evolution.
Participation in an on-call rotation and response to critical incidents outside regular working hours are part of the role.
What success looks like
Success in this role means that operational risks are identified early, capacity issues are anticipated, production changes are safer, recurring problems are permanently addressed, and the platform becomes progressively more observable, resilient, scalable, and easier to operate.
When incidents occur, they are controlled quickly, communicated clearly, and resolved with reduced recovery time. The objective is not only to respond better, but to reduce the number, frequency, and impact of incidents over time.
Plenit is a company that values operational excellence and reliability in its cloud platform. We are committed to maintaining a stable and efficient production environment while continuously improving our systems and processes.
As a Site Reliability Engineer, you will be responsible for ensuring the stability, availability, and performance of our cloud platform. You will work proactively to prevent incidents and reactively to resolve them when they occur. Your role will involve identifying risks, driving reliability improvements, and coordinating incident response efforts.
Experience with observability tools, infrastructure and system-administration tools, networking technologies, web infrastructure tools, cloud infrastructure and configuration-management platforms, Kubernetes and container orchestration technologies, incident-management and on-call tools, operational documentation tools, capacity-planning, performance-analysis, and infrastructure-scaling tools, and automation and scripting tools.
Competitive salary and benefits package.
No specific visa or eligibility requirements mentioned.
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.
Get instant access to exclusive DevOps jobs with €120K+ salaries
Best value for job search
Only €4.13/month - Save 75%