Applied ML Engineer We’re looking for an Applied ML Engineer to build systems at the intersection of machine learning research and production software.
This is an end-to-end engineering role. You should be comfortable reading a research paper, identifying what is actually testable, building the smallest useful experiment, evaluating it rigorously, and turning the result into a production system that users can interact with.
You’ll work across model evaluation, model internals, inference infrastructure, backend systems, and product interfaces. The goal is not simply to reproduce research. It is to turn promising methods into reliable, measurable, and usable products.
🏢 About SENTIENT
SENTIENT is a company focused on building systems that produce evidence people can actually trust. We are looking for someone who enjoys moving between research, experimentation, engineering, and product.
🎯 The Role
This is an end-to-end engineering role. You should be comfortable reading a research paper, identifying what is actually testable, building the smallest useful experiment, evaluating it rigorously, and turning the result into a production system that users can interact with. You’ll work across model evaluation, model internals, inference infrastructure, backend systems, and product interfaces.
✅ Key Responsibilities
Reproduce and evaluate research methods using open-weight and API-accessible models.
Work directly with model weights, logits, hidden states, activations, model APIs, and inference infrastructure when required.
Build and extend our evaluation infrastructure, including runners, judges, persistence, experiment orchestration, and reporting.
Turn research workflows into product experiences, including experiment configuration, runs, traces, comparisons, reports, and review workflows.
Investigate how verification methods behave under model modification, including fine-tuning, merging, quantization, distillation, safety removal, and deliberate evasion.
Design controlled experiments that separate meaningful signals from artifacts or confounders.
Write clear technical reports that distinguish measured evidence, interpretation, and hypotheses.
Ship production-quality systems with APIs, background jobs, observability, testing, and documentation.
📌 Required Qualifications
Strong Python engineering skills and hands-on experience with PyTorch and Hugging Face Transformers.
A strong understanding of ML evaluation, including dataset design, baselines, metrics, calibration, false positives, false negatives, statistical uncertainty, and reproducibility.
Ability to read ML research papers and implement methods from first principles rather than relying entirely on existing packages.
Experience building production software beyond notebooks, including APIs, asynchronous jobs, databases, logging, testing, and deployment.
Comfort working with open-weight models and understanding how modern LLM inference systems operate.
Ability to work across backend and frontend boundaries. Our product surface is primarily React/TypeScript, and you should be able to make complex experiments and results understandable to users.
Strong technical judgment about what experimental evidence does and does not support.
High agency and a strong sense of ownership. You are comfortable identifying problems, proposing solutions, and driving work forward without waiting for detailed instructions.
Comfortable working in a fast-moving startup environment where priorities can evolve quickly and individuals are expected to operate across functions.
⭐ Desirable Experience
Experience in any of the following is a plus:
Model provenance, fingerprinting, watermarking, distillation detection, red-teaming, safety evaluations, or interpretability.
Activation and representation analysis, probing, model hooks, logits, hidden states, or other model-internals work.
Evaluation and inference infrastructure such as DSPy, LiteLLM, Temporal, Ray, vLLM, PostgreSQL/pgvector, or similar systems.
Next.js, React, TypeScript, data visualization, or experiment dashboards.
Running and serving open-weight models on GPUs and reasoning about latency, throughput, memory, precision, and cost tradeoffs.
Designing adversarial evaluations or testing systems against deliberate attempts to evade detection.
Please let Sentient know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.