
Engineering Manager, Site Reliability
What you need
- 7+ yrs software/infrastructure/SRE experience
- 2+ yrs managing/leading engineering teams
- Bachelor's degree in CS/Engineering or equivalent experience
- Cloud infrastructure, distributed systems, containers, orchestration expertise
- Production incident response and post-incident review experience
What you'll do
- Lead and develop team of Site Reliability Engineers
- Set technical direction for reliability, scalability, performance
- Partner cross-functionally to define reliability standards
- Drive incident management and operational excellence
- Promote automation and self-service tooling to reduce toil
We're transforming the grocery industry
At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers.
Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table.
Instacart is a Flex First team
There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work.
Overview
Instacart is transforming the grocery industry by building technology that connects customers, shoppers, retailers, and brands through a reliable, convenient online marketplace. The Site Reliability Engineering team helps ensure that this experience remains resilient, scalable, and dependable as Instacart grows.
We are seeking an Engineering Manager to lead a team of Site Reliability Engineers responsible for the systems, tools, and practices that support the reliability of Instacart’s technology platform. In this role, you will manage and develop a team of engineers while partnering closely with engineering teams across the company to improve availability, scalability, performance, observability, and operational excellence.
This is an opportunity to shape the future of reliability at significant scale. You will help establish sound engineering practices, guide complex technical initiatives, and create an environment where teams can build and operate dependable systems. The role is well suited for a collaborative, hands-on people leader who is energized by complex problems, thrives in a fast-changing environment, and is motivated by helping others do their best work.
About the Job
- Lead, mentor, and develop a team of Site Reliability Engineers, establishing clear goals, providing actionable feedback, and supporting career growth and professional development.
- Set the team’s technical direction and priorities for improving the reliability, scalability, availability, performance, and operational readiness of Instacart’s systems.
- Partner with engineering, product, security, infrastructure, and other cross-functional teams to define reliability standards, influence system design, and deliver initiatives that improve the customer and developer experience.
- Drive incident management and operational excellence, including incident response, post-incident learning, service-level objectives, capacity planning, observability, and continuous risk reduction.
- Promote automation and self-service tooling that reduce operational toil, improve deployment confidence, and enable engineering teams to own and operate their services effectively.
- Balance near-term operational needs with long-term investments, making thoughtful tradeoffs in a high-growth environment where priorities and requirements can change quickly.
- Communicate clearly with technical and non-technical stakeholders, bringing transparency to reliability risks, project status, tradeoffs, and decisions.
This role requires comfort with ambiguity and a willingness to engage directly with challenging operational and organizational problems. You will be expected to make decisions with incomplete information, respond calmly during incidents, and help teams learn from failure without assigning blame. Success will require both strong people leadership and enough technical depth to ask the right questions, evaluate tradeoffs, and guide effective solutions.
About You
Minimum Qualifications
- Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
- Seven or more years of experience in software engineering, infrastructure engineering, Site Reliability Engineering, or a related field.
- Two or more years of experience managing, mentoring, or leading engineering teams.
- Professional experience with cloud infrastructure, distributed systems, networking, containers, orchestration platforms, or related production technologies.
- Experience leading or participating in production incident response, post-incident reviews, reliability improvement initiatives, and operational readiness practices.
- Experience communicating technical risks, priorities, and tradeoffs to engineering leaders and cross-functional stakeholders.
Preferred Qualifications
- Experience leading Site Reliability Engineering, platform engineering, infrastructure engineering, or developer productivity teams.
- Experience operating highly available services at significant scale and improving service-level objectives, observability, capacity, or disaster recovery capabilities.
- Experience with infrastructure as code, continuous delivery, monitoring, logging, tracing, and automated remediation.
- Experience building or evolving reliability programs across multiple engineering teams, including shared standards, operational reviews, and service ownership practices.
- Demonstrated ability to create alignment across teams, navigate ambiguity, and turn complex technical challenges into clear, achievable plans.
- A leadership approach grounded in empathy, transparency, direct communication, collaboration, and a commitment to inclusive team development.
Instacart is a remote-first organization with team members working across the United States and Canada. This role is especially well suited to candidates based on the West Coast, while applicants in other eligible locations may also be considered.
#LI-Remote
Instacart provides highly market-competitive compensation and benefits in each location where our employees work. This role is remote and the base pay range for a successful candidate is dependent on their permanent work location. Please review our Flex First remote work policy here.
Offers may vary based on many factors, such as candidate experience and skills required for the role. Additionally, this role is eligible for a new hire equity grant as well as annual refresh grants. Please read more about our benefits offerings here.
For US based candidates, the base pay ranges for a successful candidate are listed below.
Optimize your resume for this job
Get a match score and the keywords you're missing
About Instacart
Instacart is a grocery delivery startup that delivers in as short as an hour. It focuses on delivering groceries and home essentials, Instacart already has over 500,000 items from local stores in its catalogue. Customers can choose from a variety of local stores including Safeway, Whole Foods, Super Fresh, Harris Teeter, Shaw's, Mariano's, Jewel-Osco, Stanley's, and Costco. Customers can mix items from multiple stores into one order.
Similar Software roles


