We are seeking a Cloud Systems Engineer to support and operate large-scale AI and high-performance computing (HPC) environments. This role will be responsible for the deployment, maintenance, performance, and lifecycle management of GPU-accelerated compute infrastructure that powers critical AI, machine learning, and data-intensive workloads.
🏢 About Alarm.com
Alarm.com is the leading platform for intelligently connected properties. Millions of homeowners and businesses rely on Alarm.com's technology to secure, monitor, and manage their environments from anywhere. Our comprehensive suite of solutions—including security, video surveillance, access control, active shooter detection, intelligent automation, energy management, and wellness—is delivered exclusively through a trusted network of thousands of professional service providers and commercial integrators across North America and worldwide. Alarm.com's common stock is traded on Nasdaq under the ticker symbol ALRM. Alarm.com delivers serious security for serious people.
🎯 The Role
The ideal candidate is a hands-on infrastructure professional with strong Linux administration skills, deep hardware troubleshooting experience, and expertise supporting enterprise-class compute platforms. This individual will work closely with infrastructure, networking, storage, and AI engineering teams to ensure the reliability, scalability, and operational excellence of our AI infrastructure.
✅ Key Responsibilities
Deploy, configure, and maintain GPU-accelerated compute infrastructure.
Monitor system health, performance, utilization, and capacity across AI infrastructure environments.
Support infrastructure utilized for AI model training, inference, and data processing workloads.
Develop and maintain operational standards, runbooks, and maintenance procedures.
Participate in on-call support and incident response activities.
Administer enterprise Linux environments, including Ubuntu and Red Hat-based distributions.
Perform system patching, hardening, and operating system lifecycle management.
Troubleshoot operating system, kernel, storage, networking, and application-level issues.
Develop automation to streamline deployment, monitoring, and operational processes.
Support security and compliance initiatives across AI infrastructure platforms.
Install, configure, maintain, and troubleshoot enterprise compute hardware.
Diagnose and resolve issues involving GPUs, CPUs, memory, storage, power, and networking components.
Perform firmware upgrades and hardware lifecycle management activities.
Coordinate hardware replacements, vendor support engagements, and warranty services.
Participate in rack-and-stack deployments, datacenter expansions, and technology refresh projects.
Maintain accurate asset inventories and operational documentation.
Support high-performance networking technologies, including Ethernet and InfiniBand environments.
Collaborate with networking, storage, cloud, and AI engineering teams on infrastructure design and operations.
Assist with scalability, resiliency, and performance optimization initiatives.
Perform root-cause analysis of infrastructure failures and develop preventative measures.
📌 Required Qualifications
Bachelor’s degree required
3-5 years of Linux systems administration experience in production environments.
3-5 years of experience supporting enterprise server infrastructure.
Experience supporting large-scale compute environments, HPC platforms, AI infrastructure, or GPU-enabled systems.
Experience performing hardware diagnostics, firmware management, and lifecycle maintenance.
Experience working within datacenter operations environments.
Bash, Python, PowerShell, or similar scripting languages
Operating system performance tuning and monitoring
Storage and networking fundamentals
Experience with infrastructure monitoring and observability platforms, ex Grafana.
Hardware and firmware lifecycle management
🎁 Benefits
Our total rewards package is designed to support you holistically—in your health, your finances, and your life outside of work. The package includes medical plans with company subsidies, a Health Savings Account (HSA) with a company contribution, and a 401(k) with an employer match. We encourage a healthy work-life balance with paid vacation that increases with tenure, paid holidays, wellness time, and paid maternity and bonding leave. To complete the package, we also provide company-paid disability and life insurance, all within a collaborative and casual work environment.
🛂 Visa & Eligibility
Please note that sponsorship of new applicants for employment authorization, or any other immigration-related support, is not available for this position at this time.
Please let Alarm.com know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.