We are looking for a Senior DevOps Engineer / Tech Lead to join the BigBrain group. The team that builds and operates monday.com's Data Platform and leads the company's internal AI innovation. We own the infrastructure behind billions of daily events - from streaming pipelines and data orchestration to the AI Gateway and ML inference platform that power monday.com's intelligent features. We manage some of the most sensitive data in the company, operate across multiple global regions, and are responsible for keeping it all secure and running at scale.
🏢 About monday.com
monday.com is the AI work platform powering the most ambitious teams. 250,000+ customers across departments use us to bring people, workflows, and AI agents together on one flexible platform where AI doesn't just assist, it executes. We move fast, build things that matter, and foster an ownership-driven culture where you're empowered to shape how organizations work and outpace their competition.
🎯 The Role
As a Senior DevOps Engineer / Tech Lead on this team, you'll own the technical direction for critical infrastructure domains, mentor other engineers, and partner closely with product, security, data, and platform teams to build infrastructure that's secure, scalable, and increasingly autonomous. This is a high-ownership role with significant influence over the technical roadmap of our data and AI infrastructure.
✅ Key Responsibilities
Lead technical direction — Own the architecture and technical roadmap for critical infrastructure domains, and make high-impact trade-offs across teams.
Own and scale platform infrastructure — Manage multi-region Kubernetes clusters (EKS), streaming pipelines (Kafka/MSK, Debezium CDC), and data orchestration (Airflow) that handles billions of daily events.
Build and operate AI infrastructure — Deploy and maintain the AI Gateway (governance, observability, guardrails for LLM usage), ML inference platform, and the tooling that enables AI adoption across the company.
Drive infrastructure automation with AIOps — Design and build autonomous agents and intelligent tooling (n8n, LangChain, Claude Code) that automate infrastructure operations, analyze workflows, and reduce manual toil.
Drive cross-team initiatives — Lead complex, multi-team infrastructure projects end-to-end, from design through rollout.
Mentor and grow engineers — Guide and mentor other DevOps/infra engineers, review designs, and raise the technical bar across the team.
Build and maintain CI/CD & GitOps pipelines — Design and operate deployment pipelines using GitHub Actions, ArgoCD, Helm, and Terraform (CDKTF), enabling fast, safe, and reliable releases for product teams.
Ensure data security — Protect the company's most sensitive data through access control, data governance, and security-first infrastructure design.
Evolve the data platform — Work with modern data technologies (Snowflake, ClickHouse, Apache Iceberg, EMR) and contribute to the next generation of our data infrastructure.
Enable developer self-service — Improve our internal platform so engineering teams across the company can deploy, configure, and operate services independently.
Provide observability and reliability — Build monitoring, alerting, and debugging tools (Datadog, OpenTelemetry, ClickHouse) across all data and AI processes, and lead incident response for critical, high-scale systems.
📌 Required Qualifications
6-8+ years of experience as a DevOps / Infrastructure / Platform Engineer, including experience leading technical initiatives, owning architecture decisions, or mentoring other engineers.
Deep, hands-on experience with Kubernetes — cluster management, networking, scaling, and troubleshooting at production scale.
Strong experience with Infrastructure as Code (Terraform/CDKTF, Helm) and GitOps workflows (ArgoCD or similar).
Proven ownership of CI/CD pipelines — designing, maintaining, and optimizing the full release cycle for multiple teams.
Deep experience with cloud infrastructure (AWS preferred) — networking, IAM, security, cost optimization.
Security-first mindset — strong understanding of application security, access control, and data protection best practices.
Experience with data infrastructure — Kafka, Airflow, EMR, Apache Iceberg, Snowflake, ClickHouse, or similar technologies.
Fluent in Linux environments, scripting, and at least one programming language (Python, TypeScript, Go).
Proven ability to design and own system architecture at scale, and to make and defend complex technical trade-offs across multiple teams and domains.
Experience owning production reliability for critical, high-scale systems — including
⭐ Desirable Experience
Experience with AI/ML infrastructure and deployment patterns.
Experience with observability tools like Datadog, OpenTelemetry, and ClickHouse.
Experience with CI/CD tools like GitHub Actions, ArgoCD, Helm, and Terraform.
Experience with cloud platforms like AWS, GCP, or Azure.
Experience with data platforms like Snowflake, ClickHouse, Apache Iceberg, and EMR.
Please let Monday.com know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.