Software Development👥 501 employees📍 Fully Remote, Global, US
Chainlink Labs is one of the primary contributing developers of Chainlink, the backbone of blockchain. Chainlink is unifying liquidity across global markets and has enabled over $20 trillion in trans…
Job Overview
We are looking for a Senior Site Reliability Engineer (SRE) to join our Observability team at Chainlink Labs. As a Senior SRE, you will help us accelerate and enable other engineering teams by increasing self-service and decreasing cognitive load. This role requires a strong DevOps mentality, passion for building and maintaining a mature GitOps environment, and experience focusing on observability.
Key Responsibilities
Build and orchestrate a Modern OTEL-based Observability Platform
Support multiple telemetry types, like metrics, logs, and traces
Define and support modern governance in observability and problems at scale
Ensure reliability, security, and performance exceed our defined SLAs
Work with engineers from across the company to help troubleshoot issues, deploy new products and services, and increase velocity while decreasing cognitive load
Lead the design and deployment of monitoring/observability services to detect and alert the team of needed action
Ingest, aggregate, transform, and utilize data from a multitude of sources in our real-time data pipeline
Oversee the availability, performance, and supportability of our observability infrastructure
Create processes around alert response operations and support the team to ensure the reliable delivery of oracle data
Make recommendations to ensure sufficient metrics are collected to create alerts with every new feature release
Champion reliability and security by taking the time to do your work right the first time
Required Qualifications
7+ years of relevant professional experience in devops, infrastructure, SRE, and/or platform teams
Ability to develop software outside of the scope of typical infrastructure requirements and configurations
Experience programming in C, C++, Java, Python, Go, Perl, or Ruby
Expert knowledge in all aspects of designing, developing, and managing large real-time systems
Experience with monitoring and logging, including Prometheus, Grafana, and centralized logging solutions like ELK Stack, Splunk, or Grafana Stack
Experience with distributed systems and container orchestration, including maintaining or building Kubernetes clusters
Strong communication skills, including giving and receiving constructive feedback and participating in planning meetings and code reviews
About Chainlink
Chainlink is the industry-standard oracle platform bringing the capital markets onchain and powering the majority of decentralized finance (DeFi). The Chainlink stack provides the essential data, interoperability, compliance, and privacy standards needed to power advanced blockchain use cases for institutional tokenized assets, lending, payments, stablecoins, and more. Since inventing decentralized oracle networks, Chainlink has enabled tens of trillions in transaction value and now secures the vast majority of DeFi.
Many of the world’s largest financial services institutions have also adopted Chainlink’s standards and infrastructure, including Swift, Euroclear, Mastercard, Fidelity International, UBS, S&P Dow Jones Indices, FTSE Russell, WisdomTree, ANZ, and top protocols such as Aave, Lido, GMX, and many others. Chainlink leverages a novel fee model where offchain and onchain revenue from enterprise adoption is converted to LINK tokens and stored in a strategic Chainlink Reserve.
Desired Qualifications
Excitement for blockchain, Web 3.0, and similar decentralized technologies
Experience running any infrastructure in the blockchain/web3 space
Ability to scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity
Experience working remotely in a distributed team
A strong desire to grow and challenge yourself, constantly finding ways to improve and automate services to reduce toil
Tools and Services
Some of the tools and services we use daily or almost daily are: AWS; Terraform/Terragrunt; Kubernetes, Calico, and ArgoCD; Prometheus and Grafana; GitHub Actions; Packer. We expect you to be comfortable with most of those tools and very proficient in several of them.
Please let Chainlink Labs know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.