Join the team building the intelligence behind Hinge Health’s proactive member experience. The Proactive Communications & Notifications pod is creating the systems that determine what message a member receives, when they receive it, and which channel is most helpful. When these decisions are timely, relevant, and respectful, they can help members stay engaged with care and make progress toward better health outcomes.
🏢 About Hinge Health
This position will have an annual salary, plus equity and benefits. Please note the annual salary range is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, and competencies.
🎯 The Role
As a Senior AI Engineer focused on MLOps and distributed systems, you will own the production path that turns models and AI capabilities into dependable member-facing experiences. You’ll design and operate the deployment, serving, release automation, monitoring, rollback, orchestration, and backend integration systems that make AI reliable, observable, scalable, and cost-conscious in production.
✅ Key Responsibilities
Build context on Hinge Health’s member experience, communication systems, model lifecycle, data flows, and operational requirements.
Establish a clear baseline for the health of key AI-backed services, including service-level objectives, deployment paths, observability, failure modes, and operational ownership.
Ship targeted improvements to developer experience, testing, release safety, or incident response that make the team faster and more predictable.
Own the design and delivery of a production MLOps or distributed-systems initiative from technical plan through rollout, measurement, and iteration.
Build or improve safe promotion patterns for model-backed services, including versioning, configuration, feature flags, canaries, rollback, and recovery.
Partner with ML Scientists and Data Engineering to define durable interfaces, data contracts, feature inputs, freshness expectations, evaluation hooks, and production-readiness criteria.
Operate reliable online and batch inference capabilities that meet clear contracts for latency, availability, correctness, resiliency, observability, and cost.
Create reusable patterns for APIs, services, queues, durable workflows, monitoring, and incident response that help the pod integrate AI into notifications, personalization, send-time optimization, and related experiences.
Lead through influence by shaping technical decisions, mentoring engineers, raising the quality bar through design and code review, and connecting system reliability to outcomes such as message relevance, engagement, reduced communication fatigue, and sustained participation in care.
📌 Required Qualifications
3+ years of non-internship, full-time professional software engineering experience.
3+ years designing, building, and operating backend or distributed systems in production, including experience participating in on-call and incident response.
Demonstrated experience deploying and operating ML- or AI-backed services in production, with strong proficiency in Python and at least one production backend language such as TypeScript/JavaScript, Go, Java, or Kotlin.
⭐ Desirable Experience
Experience building or operating MLOps or ML platform capabilities, such as inference pipelines, model registries, deployment automation, monitoring, evaluation integration, or automated retraining.
Experience with online inference, batch scoring, recommendation systems, ranking, propensity models, or send-time optimization.
Experience with AWS and production technologies such as Kubernetes, Docker, Kafka, PostgreSQL, Airflow, Databricks, MLflow, or equivalent systems.
Experience with workflow orchestration and long-running state, including Temporal, Step Functions, Cadence, or a comparable system.
Experience building observability for ML systems, including data quality, feature freshness, model behavior, prediction quality, latency, errors, and cost.
Experience partnering with ML Scientists or Data Scientists to define production interfaces and operate models without owning model research or training.
🎁 Benefits
This position will have an annual salary, plus equity and benefits.
Please let Hingehealth know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.