You own how the platform ships and how it stays up — the reliability and security backbone behind every conversation on interface.ai. Our customers are banks and credit unions, and the bar for trust is high: real money, real regulators, and the uptime our customers count on. This role holds the platform to the enterprise-grade, 99.99% reliability we're known for and sets the engineering practices that keep it there as we scale.
🏢 About interface.ai
interface.ai is the agentic AI platform for financial services — bringing conversational and agentic AI to the credit unions and community banks that serve everyday Americans. We're not a lab, we're not a demo company, and we're not burning runway on hypotheticals. We are in production, generating real revenue, and on a mission that matters: making intelligent financial services available to the millions of people who've never had a private banker. More than 100 banks and credit unions run on interface.ai, reaching over 10 million people. Many products are live today; the biggest bets — an AI-first contact center and an AI-native consumer banking experience — are what comes next. Backed by $30M in Series A funding and already cash-flow positive, we're at the inflection point: a proven product, paying customers, and a profitable business rebuilding itself as an AI-native company to lead a world where agents, not software, do the work.
🎯 The Role
You own how the platform ships and how it stays up — the reliability and security backbone behind every conversation on interface.ai. Our customers are banks and credit unions, and the bar for trust is high: real money, real regulators, and the uptime our customers count on. This role holds the platform to the enterprise-grade, 99.99% reliability we're known for and sets the engineering practices that keep it there as we scale. You'll define the SLOs, own the deploy path, run incidents when they happen, and build the platform so a small, senior team working with AI tooling can operate it with confidence — and you'll set the reliability and security bar the rest of the org builds against.
✅ Key Responsibilities
Reliability & SLOs — define customer-facing SLIs and SLOs across product surfaces, run an error-budget program that governs release decisions, and hold the platform to its reliability bar.
Resilience & disaster recovery — own the DR strategy: regional failover, written RTO/RPO per tier, resilience against third-party dependency failure, and recovery you've actually tested rather than just planned.
Deploy & delivery — a GitOps deploy path with progressive delivery, automated analysis, and one-click rollback for every service, plus drift detection and a full change audit trail.
Cloud & infrastructure-as-code — the AWS foundation, fully managed as code; the Kubernetes platform and service mesh; capacity, cost, and scale.
Incident management — end to end: paging and severity policy, an incident-commander rotation, status-page automation, blameless post-mortems with tracked actions, and SLA reporting to customers.
Observability — metrics, logging, and distributed tracing across services; dashboards and alerts as code; burn-rate alerting that pages on what actually matters.
AI-native operations — build the automation and guardrails that let a small, senior team operate the platform with AI in the loop: runbooks encoded for safe automation, self-healing for routine work, and golden-path templates that ship new services with SLOs, alerts, secrets, and policy built in.
Security & Compliance — You own the infrastructure's security posture — including the parts unique to running an AI platform in regulated financial services.
📌 Required Qualifications
A senior-most, hands-on IC who has run production systems with real uptime commitments — and carried the pager for them.
Deep production Kubernetes on AWS, including service mesh, with GitOps-based delivery across many services.
Infrastructure-as-code at multi-account scale, including taking over and reshaping a large existing estate.
You've built an SLO and error-budget practice that actually changed release decisions, with alerting tuned to burn rate rather than noise.
You've delivered multi-region or DR capability with defined RTO/RPO and proven it with real failover tests.
Security as daily practice, not a checklist — least-privilege IAM, secrets management, admission and network policy, software supply-chain controls — and you've produced evidence that satisfied auditors (SOC 2, ISO 27001, PCI, or FFIEC-style exams).
Extreme AI fluency — you use frontier AI tools (Claude Code, Cursor) daily and have clear opinions about the boundaries agents should operate within.
Strong programming in TypeScript and/or Python, plus Bash — and writing clear enough that a regulator could follow your post-mortem.
BS/BA in Computer Science required; MS or PhD a strong plus. San Francisco-based and committed to working onsite; on-call participation. H1B transfers welcome.
⭐ Desirable Experience
Real-time voice or telephony systems, or other latency-critical streaming workloads.
Operating streaming and analytical data platforms and durable workflow engines at scale.
Please let Interface Ai know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.