We are looking for a Senior/Staff Platform Engineer to build, operate, and evolve large-scale production infrastructure. This is a hands-on platform and reliability engineering role for someone who has deep experience operating Kubernetes and cloud infrastructure, troubleshooting complex production systems, and building the automation and tooling that keeps those systems reliable.
🏢 About Virtasant
Virtasant is a company that values hands-on platform and reliability engineering. We are seeking engineers who understand the systems underneath the applications and can independently diagnose and solve infrastructure problems across multiple layers.
🎯 The Role
This is a highly autonomous role. You will work directly with technical stakeholders, own ambiguous infrastructure initiatives from design through production, and be trusted to drive technical decisions and critical issues without requiring constant direction.
✅ Key Responsibilities
Design, build, operate, and improve production Kubernetes platforms.
Own platform-level concerns including cluster architecture, networking, workload isolation, resource management, security, upgrades, scaling, and reliability.
Troubleshoot Kubernetes beyond the application layer, including networking/CNI, scheduling, node behaviour, resource constraints, controllers, and cluster-level failures.
Operate and improve large-scale, highly available infrastructure across cloud, hybrid, virtualised, and/or bare-metal environments.
Optimise platform infrastructure for reliability, performance, scalability, and operational efficiency.
Write, maintain, and improve production tooling and automation using Go, Python, or Java.
Build software and automation that improves platform operations, reliability, deployment, troubleshooting, and developer experience.
Read, debug, and contribute to existing production codebases.
Develop internal services, APIs, integrations, and operational tooling where needed.
Automate repetitive operational processes and reduce manual intervention across the platform.
Apply sound software engineering practices, including testing, code review, maintainability, and documentation.
Own the reliability and operational health of critical production infrastructure.
Lead or contribute significantly to incident response for complex platform and infrastructure issues.
Investigate root causes and implement durable remediation rather than temporary fixes.
Define and improve SLOs, SLIs, alerting, and operational processes.
Troubleshoot systems using logs, metrics, traces, profiling tools, and system-level diagnostics.
Drive improvements in availability, performance, capacity, resilience, and operational readiness.
Contribute to disaster recovery planning, testing, and continuous improvement.
Build and maintain infrastructure as code using Terraform and related automation technologies.
Create reusable infrastructure patterns and improve automation as the platform evolves.
Build and improve CI/CD and deployment workflows supporting large-scale engineering environments.
Balance delivery speed with reliability, security, scalability, and operational requirements.
Work across infrastructure provisioning, configuration management, deployment automation, and production operations.
Participate in planning and executing production cloud or infrastructure migrations, including dependency analysis, networking, cutover, rollback, and production validation.
Build and maintain production monitoring, metrics, dashboards, alerting, logging, and distributed tracing.
Improve observability so engineers can identify and diagnose failures quickly.
Use production telemetry to identify reliability, capacity, and performance problems before they become major incidents.
Continuously improve incident detection and reduce time to diagnosis and recovery.
Work directly with customer and internal engineering teams to understand requirements, investigate problems, and drive technical solutions.
Communicate architecture, technical decisions, risks, trade-offs, and progress clearly to technical stakeholders.
Own complex infrastructure initiatives from initial problem definition through design, implementation, and production operation.
Contribute to architecture discussions, RFCs, design reviews, and technical direction.
Mentor other engineers and help improve engineering and operational practices across the team.
Operate independently in ambiguous situations and take ownership when immediate technical or management direction is unavailable.
📌 Required Qualifications
10+ years of professional experience in Platform Engineering, Site Reliability Engineering, Infrastructure Engineering, DevOps, or related fields; 10+ years is preferred for Staff-level candidates.
Significant hands-on experience operating complex production infrastructure and distributed systems.
Demonstrated experience building and operating production Kubernetes platforms, not only deploying applications onto existing clusters.
Production programming experience with Go, Python, or Java.
Strong experience with production reliability, incident response, troubleshooting, and operational ownership.
Experience independently owning complex technical initiatives from an ambiguous starting point through production.
Experience working directly with technical stakeholders or customers and communicating complex technical topics effectively.
Degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
⭐ Desirable Experience
Experience planning and executing production cloud or infrastructure migrations, including cutover and rollback strategies.
Experience operating large-scale or multi-cluster Kubernetes environments.
Experience with hybrid cloud, on-premises, virtualised, or bare-metal infrastructure.
Experience building or modifying Kubernetes controllers or operators.
Advanced Kubernetes networking, CNI, NetworkPolicy, service mesh, mTLS, or workload identity experience.
Multi-cloud infrastructure experience.
Experience designing and testing disaster recovery strategies.
Experience with large-scale CI/CD or developer infrastructure.
Experience with capacity planning and performance engineering.
Experience building internal platform tooling or improving developer experience.
Experience with security, infrastructure hardening, IAM, or compliance requirements.
Previous technical leadership, mentoring, or Staff/Principal-level engineering responsibilities.
Experience working directly with external customers or stakeholders in a consulting or service-delivery environment.
🎁 Benefits
This role is currently open to candidates based in Brazil, Mexico, or Canada.
🛂 Visa & Eligibility
This role is currently open to candidates based in Brazil, Mexico, or Canada.
Please let Virtasant know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.