CloudLinux builds Linux infrastructure and security products. You will join our Automation & Management Services cell, working closely with Dmitrii Petrov to solve problems across teams and services: cloud-cost data, infrastructure inventory, network policies and capacity workflows.
🏢 About CloudLinux
Check out our website for more information https://cloudlinux.com/
🎯 The Role
Inside the Infrastructure Department, the Platform cell is a small team. We run the observability platform, the company's GitLab and the CI runners behind it, a few smaller engineering services, and the automation the department relies on for provisioning and configuration.
✅ Key Responsibilities
Run the observability platform. Keep it healthy, onboard teams, watch cost and capacity, and maintain the alerting that runs on top of it.
Run GitLab and the CI runner fleet. Upgrades, capacity, access, backups and restore drills.
Keep the rest of our services healthy, with the monitoring and runbooks a production service needs.
Deploy new services when they are requested. Research the options, pick a design, and stand the service up from scratch according to good practice: as code, monitored, backed up, documented.
Work with developers' requests. Access, onboarding, pipeline problems, new exporters and dashboards. Answer them, and turn the recurring ones into self-service.
Run incidents. Diagnose and mitigate impact, restore service safely, then complete the root-cause analysis and post-mortem. Deliver the prevention or detection improvements the incident calls for.
Ship everything as code, reviewed in merge requests. Plan and check before every change.
Write for engineers outside the team. Runbooks, onboarding guides, maintenance notices and status updates that people can act on.
Work with AI agents. Delegate collection and drafting to them, review their output as you would a colleague's merge request, and record what you learn where the team can find it.
📌 Required Qualifications
Senior-level experience in infrastructure, platform or site reliability engineering, including at least one production service you were responsible for keeping up.
Linux systems administration and debugging on bare metal and virtual machines. Much of our infrastructure is not Kubernetes.
Kubernetes in production delivered through GitOps, including cluster upgrades you performed yourself.
Infrastructure as code as your delivery form: Ansible and Terraform or OpenTofu, changes reviewed in merge requests.
GitLab administration and GitLab CI in production, self-hosted or SaaS. Deep experience with another CI system is acceptable if you can show the same depth.
Working knowledge of the Prometheus and Grafana ecosystem: you have run it for a team, written alert rules and dashboards, and can read PromQL.
Written technical explanation for engineers outside your team: runbooks, notices, answers to requests.
Strong communication and interpersonal skills. This role deals with people at least as much as with servers: most work starts as a conversation with a product team, and you need to understand what they actually need, agree scope, priority and timing with them, push back politely when a request should not be done as asked, and keep everyone informed while the work is in progress. We are looking for someone other teams enjoy working with.
Advanced use of AI engineering assistants such as Claude and Codex: providing context, breaking down tasks, designing agent loops, and delegating plans for unattended, end-to-end execution within defined scope and permissions, with clear stop conditions. You can explain, debug and test the resulting automation, and verify generated commands, scripts and conclusions before they touch production.
English - upper-intermediate or higher - to ensure clear communication of progress within the teams.
⭐ Desirable Experience
Alerting design: SLOs, burn-rate alerts, thresholds sized from data.
MicroVM isolation for CI: Kata Containers, Firecracker or gVisor.
S3-compatible object storage operations: Ceph RGW or similar.
AWS with real cost work.
Self-hosted Sentry, or another Kafka, ClickHouse and Redis-backed application you have kept alive under load.
Python or Go for exporters and small internal services.
Please let Cloud Linux know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.