Own the reliability, scalability, and performance of Peec AI’s core systems and infrastructure. Design, build, and maintain the tooling, automation, and monitoring that keep our services fast, secure, and highly available.
Key Responsibilities
Partner closely with product and engineering teams to ensure new features are reliable, observable, and easy to operate from day one
Develop and refine incident response practices, ensuring issues are triaged quickly and resolved with minimal user impact
Proactively identify and address bottlenecks, single points of failure, and operational inefficiencies across the stack
Champion operational excellence and a culture of reliability, driving best practices across the engineering organization
Required Qualifications
5+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or similar roles supporting production systems at scale
Deep expertise with Infrastructure as Code tools (Terraform, Pulumi, CloudFormation, etc.)
Strong experience with observability platforms (e.g., Datadog, Sentry, Prometheus, Grafana) and incident response tooling (PagerDuty, Incident.io, or similar)
Proven proficiency with major cloud platforms (GCP, AWS, or Azure) and modern distributed systems
Strong programming and scripting skills (e.g., TypeScript and Python) for automation and tooling
Solid understanding of CI/CD, Kubernetes, containerization, networking, databases, and cloud security principles
Excellent problem-solving skills, attention to detail, and a strong commitment to operational excellence
Bonus Points
Experience supporting AI/ML workloads or data-intensive systems
Prior SRE experience in a high-growth startup or globally distributed infrastructure environment
Familiarity with zero-downtime migrations, multi-region architectures, or compliance frameworks
Please let Peec AI know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.