Join the Platform R&D Group, where we leverage data to drive real impact in the maritime world. Our core mission is to become the intelligence authority of the oceans - mission-grade maritime AI that brings action-ready clarity across the global seas. We are scaling our operations, rapidly advancing our AI efforts, models, and the complex problems we are solving. For that, we are building robust infrastructure that will enable us to push more models to production, in a faster manner and with our eyes wide open on how our models perform.
Key Responsibilities
Lead and drive the deployment, lifecycle management, and monitoring of ML/DL models in all stages leading to production.
Design and implement systems for Dataset and Label Management, including versioning and integrating customer feedback into labeling workflows.
Establish and maintain a robust Model Repository/Registry that supports versioning, local inference, and model lineage.
Lead the implementation of advanced Experiment Tracking and Monitoring solutions for both Data Science and Generative AI, focusing on evaluation, data drift detection, and model reproducibility.
Own model serving and inference systems—including autoscaling, A/B testing, canary rollouts, and latency/cost optimization for production models.
Enable specialized infrastructure for Generative AI capabilities, including tagging tools, prompt management, and LLM testing services.
Drive operational excellence by improving tool deployment usability and implementing granular cost visibility across projects and environments.
Developing reusable components such as standardized data loaders, CI/CD pipelines, and automated workflows for tasks like model retraining.
Collaborate directly with Data Scientists and the rest of the Data Platform Engineering team to productionize ML/DL models developed for cloud environments.
Required Qualifications
B.Sc. or M.Sc. in Computer Science or Software Engineering or related field
Experienced with ML/DL workflows and their best practices
Experienced with CI/CD workflows and their best practices
Worked with public cloud (AWS/Azure/GCP)
Experienced with Python and Java
Experience with various data stores like Postgres, MongoDB, Redis
Experience with DS tools such as MLFlow, Langfuse, SageMaker, etc.
Experience with Spark
Experience with PyTorch/TensorFlow
Benefits
Flexible work arrangements.
15 working days per year as Non-Operational Allowance for personal recreation, fully compensated.
Please let Globaldevgroup know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.