We’re the experience innovation company - a trusted partner to the world’s most recognized brands. To our people we offer growth opportunities, values-driven culture, international careers and the chance to shape the future of experience.
🏢 About Valtech
Valtech is the experience innovation company that exists to unlock a better way to experience the world. By blending crafts, categories, and cultures, we help brands unlock new value in an increasingly digital world.
At the intersection of data, AI, creativity, and technology, we drive transformation for leading organizations, including L’Oréal, Mars, Audi, P&G, Volkswagen Dolby, and more.
At Valtech, we don’t just talk about transformation. We make it happen. Our people are the heart of our success, and we foster a workplace where everyone has the support to thrive, grow and innovate.
🎯 The Role
At Valtech, you’ll find an environment designed for continuous learning, meaningful impact, and professional growth. Whether you're pioneering new digital solutions, challenging conventional thinking or building the next generation of customer experiences, your work will help transform industries.
We are proud of:
The work we do and the innovation we drive
Our values of share, care and dare
A workplace culture that fosters creativity, diversity and autonomy
Our borderless, global framework, which enables seamless collaboration
✅ Key Responsibilities
Monitor cloud infrastructure (AWS, Azure, and/or GCP) using observability tooling; triage alerts and respond to incidents within defined SLAs as part of a 24x7 rotation, including nights, weekends, and holidays as scheduled.
Lead incident response, root cause analysis, and problem management for production issues, escalating to vendors or engineering leadership when required.
Execute and support change requests, patching, backups, and routine maintenance following Valtech's change management and least-privilege access processes.
Design, build, and maintain Infrastructure as Code (e.g., Terraform, CloudFormation, ARM/Bicep) and CI/CD pipelines to automate provisioning and reduce manual toil.
Own and continuously improve monitoring, alerting, and dashboarding (e.g., CloudWatch, Datadog, Grafana, Azure Monitor) to increase visibility and reduce mean time to detect/resolve.
Support security operations: apply patches, respond to vulnerability findings, and follow incident escalation procedures in line with client and Valtech security policies.
Maintain accurate runbooks, knowledge base articles, and shift handover notes to ensure smooth 24x7 continuity across teams and time zones.
Collaborate with client stakeholders, account teams, and cross-functional Valtech engineering teams on service improvement initiatives.
Contribute to capacity planning, cost optimization, and reliability improvements (SLOs/SLIs) for managed environments.
Participate in on-call rotation and follow escalation/paging procedures for priority incidents outside scheduled shifts, where applicable.
📌 Required Qualifications
Advanced/Fluent English
5+ years of experience in cloud engineering, DevOps, SRE, or infrastructure support roles, ideally within a managed services or 24x7 operations environment.
Deep, hands-on experience with at least one major cloud platform (AWS, Azure, or GCP); multi-cloud experience is preferred.
Working knowledge of Linux and/or Windows Server administration.
Scripting/automation experience (e.g., Python, Bash, PowerShell) and familiarity with Infrastructure as Code tools.
Solid experience with containerization and orchestration (Docker, Kubernetes) is expected.
Familiarity with monitoring and incident management tooling (e.g., Datadog, PagerDuty, ServiceNow, Jira Service Management).
Understanding of ITIL-aligned incident, problem, and change management practices.
Strong troubleshooting skills and the ability to remain calm and methodical during high-pressure incidents.
Clear written and verbal communication skills for shift handovers and client-facing updates.
Willingness and ability to work rotating shifts, including nights, weekends, and public holidays, to support 24x7 coverage.
High adaptability and curiosity around emerging tools
Strong analytical judgement to assess AI outputs
Systems thinking with the ability to understand full incident to resolution chains with AI support
Experience with an AI-first approach to triage, RCA generation and observability insights
⭐ Desirable Experience
Cloud certification (e.g., AWS Certified Solutions Architect Professional/DevOps Engineer Professional, Microsoft Azure Administrator/Solutions Architect Expert, Google Cloud Professional Cloud Architect).
Experience supporting enterprise or multi-tenant client environments in a managed services or MSP setting.
Exposure to security frameworks and compliance requirements (e.g., ISO 27001, SOC 2, GDPR).
Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
Experience with IT Service Management platforms enriched by AI
🎁 Benefits
This is a Full Time position based in Brazil, open for remote or hybrid work (São Paulo or Florianópolis).
Beyond a competitive compensation package, we offer:
Flexibility, with remote and hybrid work options (country-dependent)
Career advancement, with international mobility and professional development programs
Learning and development, with access to cutting-edge tools, training and industry experts
🛂 Visa & Eligibility
We are committed to inclusion and accessibility. If you need reasonable accommodations during the interview process, please either indicate it in your application or let your Talent Partner know.
Please let Valtech know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.