We're transforming the grocery industry. At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table.
🏢 About Instacart
Instacart is transforming the grocery industry by building technology that connects customers, shoppers, retailers, and brands through a reliable, convenient online marketplace. The Site Reliability Engineering team helps ensure that this experience remains resilient, scalable, and dependable as Instacart grows.
🎯 The Role
We are seeking an Engineering Manager to lead a team of Site Reliability Engineers responsible for the systems, tools, and practices that support the reliability of Instacart’s technology platform. In this role, you will manage and develop a team of engineers while partnering closely with engineering teams across the company to improve availability, scalability, performance, observability, and operational excellence.
✅ Key Responsibilities
Lead, mentor, and develop a team of Site Reliability Engineers, establishing clear goals, providing actionable feedback, and supporting career growth and professional development.
Set the team’s technical direction and priorities for improving the reliability, scalability, availability, performance, and operational readiness of Instacart’s systems.
Partner with engineering, product, security, infrastructure, and other cross-functional teams to define reliability standards, influence system design, and deliver initiatives that improve the customer and developer experience.
Drive incident management and operational excellence, including incident response, post-incident learning, service-level objectives, capacity planning, observability, and continuous risk reduction.
Promote automation and self-service tooling that reduce operational toil, improve deployment confidence, and enable engineering teams to own and operate their services effectively.
Balance near-term operational needs with long-term investments, making thoughtful tradeoffs in a high-growth environment where priorities and requirements can change quickly.
Communicate clearly with technical and non-technical stakeholders, bringing transparency to reliability risks, project status, tradeoffs, and decisions.
📌 Required Qualifications
Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
Seven or more years of experience in software engineering, infrastructure engineering, Site Reliability Engineering, or a related field.
Two or more years of experience managing, mentoring, or leading engineering teams.
Professional experience with cloud infrastructure, distributed systems, networking, containers, orchestration platforms, or related production technologies.
Experience leading or participating in production incident response, post-incident reviews, reliability improvement initiatives, and operational readiness practices.
Experience communicating technical risks, priorities, and tradeoffs to engineering leaders and cross-functional stakeholders.
⭐ Desirable Experience
Experience leading Site Reliability Engineering, platform engineering, infrastructure engineering, or developer productivity teams.
Experience operating highly available services at significant scale and improving service-level objectives, observability, capacity, or disaster recovery capabilities.
Experience with infrastructure as code, continuous delivery, monitoring, logging, tracing, and automated remediation.
Experience building or evolving reliability programs across multiple engineering teams, including shared standards, operational reviews, and service ownership practices.
Demonstrated ability to create alignment across teams, navigate ambiguity, and turn complex technical challenges into clear, achievable plans.
A leadership approach grounded in empathy, transparency, direct communication, collaboration, and a commitment to inclusive team development.
🎁 Benefits
Instacart provides highly market-competitive compensation and benefits in each location where our employees work. This role is remote and the base pay range for a successful candidate is dependent on their permanent work location. Offers may vary based on many factors, such as candidate experience and skills required for the role. Additionally, this role is eligible for a new hire equity grant as well as annual refresh grants.
🛂 Visa & Eligibility
Instacart is a remote-first organization with team members working across the United States and Canada. This role is especially well suited to candidates based on the West Coast, while applicants in other eligible locations may also be considered. Currently, we are only hiring in the following provinces: Ontario, Alberta, British Columbia, and Nova Scotia.
Please let Instacart know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.