Shield AI is a venture-backed defense-tech company with the mission of protecting service members and civilians with intelligent systems. We are looking for a Staff Platform Engineer to contribute to the Platform solutions and infrastructure that power Forge, the AI Factory. These components provide distributed runtime capabilities that enable teams to reliably orchestrate workloads, process data, and underpin the workflows of building an AI Pilot.
🏢 About Shield AI
Shield AI is a venture-backed defense-tech company with the mission of protecting service members and civilians with intelligent systems. Its products include Hivemind autonomy software, V-BAT and X-BAT aircraft, and Aechelon simulation and synthetic reality technologies. With offices and facilities across the U.S., Europe, the Middle East, and Asia-Pacific, Shield AI’s technology actively supports operations worldwide.
🎯 The Role
This is a hands-on technical leadership role. You will define platform architecture, implement production software, establish reusable operational patterns, and partner with teams across autonomy, ML, simulation, test, infrastructure, and product engineering. Success requires balancing developer productivity, reliability, extensibility, performance, portability, and long-term operational maintainability.
✅ Key Responsibilities
Build Kubernetes-native platform services: Develop and operate Kubernetes-based services, controllers, operators, deployment patterns, and runtime integrations that support distributed workloads across multiple environments.
Develop distributed orchestration capabilities: Design and build reusable primitives for authoring, scheduling, and scaling pipeline work.
Build reliable data-processing infrastructure: Develop platform capabilities for data storage, ingestion, validation, transformation, and governance.
Develop highly extensible platform components: The Forge Platform base provides standardized tooling around authentication, authorization, observation, networking, routing, secret management, and more to the services that are integrated on top of the ecosystem.
Create reference architectures: Establish recommended deployment patterns, operating profiles, capacity guidance, benchmarks, reliability practices, and distribution approaches across cloud providers, on-prem, edge, and air-gapped environments.
Advance observability and operability: Establish end-to-end metrics, logs, traces, structured events, dashboards, alerting, service-level objectives, operational diagnostics, and runbooks for workflows, pipelines, event streams, and platform services.
Partner with downstream teams: Work directly with autonomy, ML Ops, simulation, test, infrastructure, product, and customer-facing teams to turn recurring distributed-systems problems into reusable platform capabilities.
📌 Required Qualifications
Significant experience designing and operating production distributed systems, cloud-native platforms, backend infrastructure, or data-intensive services.
Strong software engineering skills and a record of delivering production systems in Go and Python.
Deep understanding of distributed-systems fundamentals, including failure handling, idempotency, consistency tradeoffs, retries, ordering, delivery semantics, backpressure, partitioning, state management, and fault tolerance.
Experience designing or operating workflow orchestration, distributed job execution, or data processing pipelines.
⭐ Desirable Experience
No specific desirable experience mentioned.
🎁 Benefits
Competitive salary and benefits package.
🛂 Visa & Eligibility
No specific visa or eligibility information mentioned.
Please let Shield AI know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.