You'll own the GPU infrastructure Luma's research and product run on — thousands of NVIDIA and AMD GPUs across on-prem and multi-cloud (AWS and OCI). As a Senior SRE, you keep training and inference clusters reliable and fast, and you help redesign them for the next level of scale. This is a hands-on, close-to-the-metal role for a first-principles Linux engineer. You'll be the final escalation for the hardest GPU, networking, and kernel-level failures, sometimes debugging directly with NVIDIA. It fits someone who thrives on low-level problems in a fast, less-structured environment. If you want a narrow, well-bounded ops role, this isn't it.
🏢 About Luma
Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.
🎯 The Role
Team: Infra Reliability · SF Bay Area / Remote (US)
✅ Key Responsibilities
Take end-to-end ownership of production GPU clusters for training and inference across AWS and OCI, keeping them highly available and performant.
Join critical re-architecture sessions to redesign systems for higher efficiency and scale.
Tune Linux performance deeply, at the OS and kernel level.
Build automation in Python, Go, or Bash to manage, monitor, and self-heal infrastructure without heavy toil.
Serve as the final escalation for the hardest GPU, networking (InfiniBand/RDMA), and system failures, working with vendors like NVIDIA.
Help achieve and maintain security certifications (SOC 2 Type 1 & 2, ISO) with strong infrastructure security practices.
📌 Required Qualifications
5+ years as an SRE, production, or infrastructure engineer in a fast-paced, large-scale environment.
Deep, hands-on Linux expertise, containerized systems, and low-level performance debugging.
Working experience with Terraform, Airflow, and Ray.
Strong experience with AWS or OCI.
Practical experience with high-performance networking (InfiniBand, RDMA, or RoCE).
Working knowledge of security best practices and compliance frameworks like SOC 2 and ISO.
Comfort in a less-structured, fast-paced environment.
⭐ Desirable Experience
Deep expertise with GPU tooling for NVIDIA and AMD (DCGM, ROCm).
Experience managing large-scale GPU clusters for AI/ML training or inference.
Familiarity with Kubernetes or orchestration frameworks like Ray.
Deep expertise in data pipelines and infrastructure.
Please let Luma AI know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.