As an SRE in Vehicle Software, you will keep Wayveβs autonomous driving fleet reliable, observable, and safe while it operates on public roads. You will work at the boundary of software, hardware, and operations, turning real-world incidents and performance bottlenecks into lasting engineering improvements. This role offers a direct line of sight from what you build to safer deployments, faster iteration, and greater fleet scale.
π’ About Wayve
Wayve is an AI company building the future of autonomous driving. We are a team of engineers, researchers, and product managers working together to create safe and reliable AI for vehicles.
π― The Role
You will be responsible for ensuring the reliability, availability, and performance of vehicle software systems used across the dev fleet. This involves taking part in a team on-call rotation, building and operating monitoring, logging, alerting, and on-call tooling, and driving incident response and post-incident learning.
β Key Responsibilities
Own and improve the reliability, availability, and performance of vehicle software systems used across the dev fleet.
Take part in a team on-call rotation, providing out-of-hours support for live systems when required.
Build and operate monitoring, logging, alerting, and on-call tooling that enables fast detection, diagnosis, and recovery.
Drive incident response and post-incident learning, translating root causes into durable fixes and preventative controls.
Design and deliver automation for fleet operations, deployments, and repetitive workflows to reduce manual intervention.
Partner closely with Vehicle SW, operations, and platform teams to define SLOs, reliability metrics, and release readiness.
Continuously harden the production environment through capacity planning, change management, and reliability-focused reviews.
π Required Qualifications
Proven experience in an SRE, production reliability, or platform operations role for complex distributed systems.
Strong Linux fundamentals and hands-on experience with CI/CD, containers (Docker), and orchestration (Kubernetes).
Proficiency in at least one systems or scripting language (Python, C++, or Rust) with a bias for automation.
Deep troubleshooting skills across networking, distributed systems, and databases, including performance and availability issues.
Experience designing observability stacks and using tools such as Datadog, Prometheus, Grafana, OpenTelemetry, Splunk, or Humio.
Clear communication skills, including incident leadership, writing postmortems, and influencing engineering priorities.
β Desirable Experience
Cloud platform experience (AWS, GCP, or Azure), including infrastructure-as-code and secure production operations.
Experience with real-time or safety-critical systems, hardware-in-the-loop, or embedded/robotics environments.
Familiarity with fleet operations, telemetry pipelines, and operating software on edge devices at scale.
Experience defining and running SLOs/SLIs and reliability programs across multiple teams.
π Benefits
Hybrid working policy: 3 days a week minimum in the office in Germany, Baden-WΓΌrttemberg, with the rest of the time spent working from home.
π Visa & Eligibility
No specific visa or eligibility information provided.