Software Development👥 201 employees📍 San Francisco, California, USEst. 2012
Lambda provides computation to accelerate human progress. We're a team of Deep Learning engineers building the world's best GPU cloud, clusters, servers, and workstations. Our products power engineer…
Job Overview
Senior Site Reliability Engineer - Observability position at Lambda. Join us in building the world's best deep learning cloud.
Key Responsibilities
Deploy and operate observability platforms for logging, metrics, and distributed tracing
Automate deployment and operation of observability systems
Set up monitoring for modern AI/HPC clusters
Develop platform software to improve system reliability
Lead engineering teams to design monitoring solutions
Required Qualifications
8+ years software engineering experience, 3+ years in Go
5+ years Site Reliability Engineering experience
Understanding of Observability tools and practices
Experience with Kubernetes, CI/CD pipelines
Collaborative and quality-focused mindset
About Lambda
Founded in 2012, Lambda is growing fast with generous compensation and notable investors. We offer health, dental, vision coverage, wellness stipends, 401k plan, and flexible paid time off.