Spare is a fast-growing, successful startup. We thrive on innovation, rapid execution, and delivering top-tier products that make a real impact in the transit industry. We are now seeking a Software Development Team Lead to join our Infrastructure Platform Team and help build and operate the cloud platform that every Spare product runs on.
🏢 About Spare
Spare is a fast-growing, successful startup. We thrive on innovation, rapid execution, and delivering top-tier products that make a real impact in the transit industry.
🎯 The Role
As leader of the team responsible for Spare's core infrastructure platform, you will play a critical role in designing and building the reliable, secure, and cost-efficient foundation that powers our customers' on-demand transit systems. Working closely with others on the team, you will own the correct and efficient operation of the platform across a variety of real-world production scenarios.
In this role, you'll split your time 50/50 between hands-on technical contribution and people leadership, working in an autonomous environment where you'll solve interesting technical challenges while growing and mentoring a high-performing team. Currently there are four software developers reporting to this role.
Given the nature of our business, this role requires someone who can balance technical expertise with strong product sensibilities, creating solutions that are both technically sound and accessible to end users. The role includes some travel as part of the job responsibilities – specifically, up to four customer site visits per year to gain firsthand insights, plus participation in our biannual software development hackathons in Vancouver.
✅ Key Responsibilities
Own the design and development of core infrastructure platform capabilities from inception to launch
Build and evolve tooling, automation, and platform services that make every engineering team at Spare faster and safer
Architect and implement high-performance, scalable distributed systems on GCP and Kubernetes
Drive improvements in cluster reliability, application resilience, and internal access security
Operate and maintain Spare's Redis and PostgreSQL databases — availability, performance tuning, scaling, upgrades, backups, and disaster recovery
Drive Spare toward AI SRE practices — embed AI agents into incident detection, alert triage, and operational workflows to make reliability work faster and more proactive
Manage and continuously improve the SRE on-call rotation — healthy schedules, clear escalation paths, and blameless post-mortems that turn incidents into systemic improvements
Drive FinOps practices across the organization — own cloud spend visibility, right-sizing, and cost optimization initiatives that deliver measurable savings
Use AI agentic tooling daily to accelerate your own and your team's engineering output, and coach the team to do the same
Actively mentor software developers of all levels and uplift team capacity
Collaborate cross-functionally with product managers, designers, and other software developers
Ensure 99.99% uptime and maintain exceptional system performance
Participate in team agile rituals and help improve software development processes
📌 Required Qualifications
7+ years of software development experience, with at least 2+ years in a people leadership role
Expert in backend technologies with strong distributed systems experience
Demonstrated proficiency with AI-assisted and agentic development workflows (AI coding agents, automation of engineering and operational tasks)
Experience operating systems at scale with a strong reliability and uptime mindset
Experience running or managing an SRE on-call rotation, including incident response and post-mortem culture
Deep experience with cloud infrastructure (GCP) and container orchestration (Kubernetes)
Experience driving cloud cost optimization and FinOps initiatives — right-sizing, spend visibility, and measurable cost savings
Experience with infrastructure-as-code and configuration management tooling (Terraform)
Understanding of security best practices, especially around internal access control and container security
Demonstrated success in managing a team of software developers, with a focus on team and individual performance
Demonstrated ability to mentor other developers and provide technical leadership
Strong problem-solving, debugging, and system design skills
Excellent communication and collaboration skills
⭐ Desirable Experience
Experience in the transit industry or another real-time, safety-critical domain
Experience building internal developer platforms and golden-path tooling
Experience applying AI/LLM tooling to SRE or infrastructure operations (AIOps, automated incident triage)
Experience with CI/CD systems at scale, including test sharding and build performance
Experience operating and maintaining PostgreSQL and Redis in production — performance tuning, indexing, replication, backup and restore
🎁 Benefits
Compensation: CA$150K – CA$260K Equity offered as a variable dependent on base salary.