Elicit is an AI research assistant that uses language models to help researchers figure out what’s true and make better decisions, starting with common research tasks like literature review. We're looking for an Infrastructure Engineer to own and evolve our infrastructure platform that underpins Elicit's product, ensuring it is reliable, secure, cost-efficient, and ready for the demands of a growing enterprise customer base.
🏢 About Elicit
Elicit is an AI research platform used by scientists, pharma companies, and decision-makers for high-stakes evidence synthesis. Our mission is to radically increase the amount of good reasoning in the world. We aim to push the frontier forward for experts and make good reasoning more affordable for non-experts. Elicit is a scalable ML system based on human-understandable task decompositions, with supervision of process, not outcomes.
🎯 The Role
This is the first dedicated infrastructure hire at Elicit. You will own our cloud infrastructure across AWS and GCP, scale single-tenant deployments, build our observability and incident response practice, drive compliance and security operations, manage infrastructure cost and capacity, improve developer experience, and contribute to backend systems where infrastructure and application intersect.
✅ Key Responsibilities
Own our cloud infrastructure across AWS and GCP — Kubernetes clusters, networking, databases (Aurora PostgreSQL, Redis, MongoDB Atlas), Cloudflare, and our CI/CD pipeline.
Scale single-tenant deployments from a handful to many — each with distinct data retention, geographic, monitoring, and compliance requirements.
Build our observability and incident response practice — proactive monitoring, alerting, SLA tracking, and structured post-mortems.
Drive compliance and security operations — ensure we follow through on the policies we've written (SOC 2, NIST AI framework, EU Cyber Resilience).
Manage infrastructure cost and capacity — make smart decisions about where we run workloads (AWS, CoreWeave, Parasail), optimize spend, and plan capacity as usage grows.
Improve developer experience — CI/CD pipeline performance, preview environments, local development tooling, and deployment confidence.
Contribute to backend systems where infrastructure and application intersect — circuit breakers, inference routing, data connector infrastructure for enterprise customers bringing their own data.
📌 Required Qualifications
5+ years of hands-on infrastructure/SRE/platform engineering experience.
An AI-native way of working. Agentic coding tools (Claude Code, Cursor, Devin, etc.) are how we build at Elicit, and infrastructure is no exception: agents help us write IaC and investigate incidents.
Solid Terraform experience. This is our primary infrastructure-as-code layer and the most important technical requirement.
Strong Kubernetes expertise. You've operated production clusters, not just deployed to them. Comfortable with EKS, networking, autoscaling (Karpenter), and debugging cluster-level issues.
AWS experience (primary), with GCP familiarity a plus.
GitOps and CI/CD fluency. Argo CD, GitHub Actions, or equivalent. You understand deployment automation, rollback strategies, and change management.
SRE mindset. You've built or significantly improved observability stacks (DataDog or equivalent), incident response processes, and on-call practices.
Security and compliance awareness. Experience with SOC 2 or similar frameworks, SIEM tooling, and translating compliance requirements into technical controls.
⭐ Desirable Experience
You've written about, spoken about, or built projects demonstrating your AI-native way of working.