We are looking for an MLOps Engineer to build, scale, and operate the critical systems that power Sequen’s AI models in production. This is a foundational, purely infrastructure-focused role sitting at the intersection of machine learning, backend distributed systems, and platform performance.
🏢 About Sequen
Sequen provides an integrated platform that pairs cutting-edge frontier ranking models with the infrastructure to run them in production—at sub-10ms latency and enterprise scale. The world's largest retailers, marketplaces, and travel platforms use Sequen to rank, recommend, and personalize, with an autonomous research engine that compounds model performance into revenue and margin lift measured in hundreds of millions of dollars per customer. We are a small, highly technical, early-stage team focused on turning recent advances in AI into production-grade systems that operate under unforgiving real-world constraints.
🎯 The Role
You will not be client-facing; instead, your primary customer will be our internal ML research scientists. Your mission is to make model serving, evaluation, and scaling completely seamless, reliable, and highly optimized in high-throughput production environments.
✅ Key Responsibilities
Build ML infrastructure: Design, operate, and maintain robust systems for low-latency model deployment, distributed inference pipelines, and automated real-time telemetry.
Scale ranking systems: Move models cleanly from experimentation to production, optimizing the critical trade-offs between execution latency, GPU/CPU throughput, and cloud infrastructure costs.
Implement model CI/CD: Build reliable infrastructure for automated model versioning, canary releases, hot-swappable container rollouts, and zero-downtime rollbacks.
Drive system observability: Architect and monitor real-time pipelines to track model performance, data distribution drift, and system reliability anomalies.
Develop evaluation loops: Engineer robust evaluation pipelines and feedback loops to continuously validate live inference accuracy and prevent training-serving skew.
Optimize platform bottlenecks: Proactively isolate and eliminate performance bottlenecks across our serving layers, improving core tooling, model warm-up times, and researcher velocity.
Collaborate with research: Partner closely with our internal ML researchers and backend engineers to translate experimental model breakthroughs into resilient, production-grade serving topologies.
📌 Required Qualifications
Proven track record: Bring 4–8+ years of practical experience in MLOps, Machine Learning Engineering, or distributed platform/infrastructure engineering.
Low-latency serving expertise: Demonstrate hands-on experience deploying and serving ultra-low-latency machine learning models under heavy, real-time concurrent workloads.
Core ML framework mastery: Maintain deep, production-grade proficiency with Python and PyTorch.
Cloud & container fluency: Operate comfortably across major cloud platforms (AWS, GCP, or Azure) utilizing modern containerization and orchestration tooling (Docker, Kubernetes).
Pipeline engineering depth: Show experience designing robust, scalable data pipelines, model registries (e.g., MLflow), and automated CI/CD infrastructures.
Systems core maturity: Bring a solid, first-principles understanding of the complete machine learning lifecycle, asynchronous event-driven patterns, and distributed systems.
⭐ Desirable Experience
Rust systems proficiency: Bring production experience or active, hands-on familiarity with Rust for low-overhead systems engineering.
Generative AI experience: Exposure to serving and optimizing large language models (LLMs) or large-scale generative model architectures (vLLM, Triton).
Modern MLOps tooling: Familiarity with enterprise-grade feature stores, advanced experiment tracking, and systematic model evaluation frameworks.
Startup velocity: Prior experience building and scaling software infrastructure from scratch in fast-moving, early-stage, or hypergrowth startups.
🎁 Benefits
High-impact influence: A foundational, high-autonomy role directly shaping the core deployment and serving topology of a category-defining AI infrastructure company.
Pioneering systems: The unique opportunity to build and scale category-defining, low-latency ML platforms backed by proven, highly quantified customer revenue results.
Top-tier reward: Highly competitive base salary, uncapped performance metrics, and meaningful early-employee equity.
Premium benefits: Full premium medical/dental/vision coverage, unlimited paid time off, and a highly collaborative, world-class environment.
Please let Sequen Ai know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.