Harness is the AI Software Delivery Platform company, led by technologist and entrepreneur Jyoti Bansal (founder of AppDynamics, acquired by Cisco for $3.7B). Harness has raised approximately $570M in funding and is valued at $5.5B, backed by leading investors including Goldman Sachs, Menlo Ventures, IVP, Unusual Ventures, Citi Ventures, and more. As AI accelerates code creation, the real bottleneck has shifted to everything after the code – testing, deployments, application security, reliability, compliance, and cost optimization. Harness brings AI and automation to this “outer loop,” helping teams ship software faster while maintaining security and governance throughout the entire software delivery lifecycle.
🏢 About Harness
Powered by Harness AI and the Software Delivery Knowledge Graph, the Harness Platform applies deep context and intelligent automation across the software delivery lifecycle with governance and policy-driven controls embedded throughout the platform.
🎯 The Role
As a Staff Software Engineer on the AI-SRE team, you will design and build intelligent, scalable platforms that improve service reliability, incident response, and operational efficiency. You will provide technical leadership across the team, own critical components, influence architecture, and collaborate with Site Reliability Engineers and cross-functional teams to solve complex production challenges.
✅ Key Responsibilities
Design, develop, and maintain scalable, highly available AI-SRE platform services.
Define technical architecture and author functional specifications and design documents.
Own critical system components from design through production operation.
Diagnose complex issues across distributed systems and production environments, particularly involving multi-source event processing.
Build AI-assisted capabilities for incident detection, diagnosis, remediation, and automation.
Establish engineering standards for quality, scalability, security, performance, and reliability.
Identify technical debt and scaling risks, then drive improvements across the platform.
Design and develop REST, gRPC, GraphQL, and event-driven APIs.
Partner with SRE, platform, product, and infrastructure teams during incident investigation.
Define observability strategies using metrics, logs, traces, and actionable alerts.
Lead technical reviews of architecture, specifications, designs, and code.
Mentor Software Engineers and Senior Software Engineers through design reviews and architectural guidance.
Influence technical direction across teams while balancing delivery and long-term maintainability.
Evaluate emerging AI and platform technologies and apply them to practical reliability problems.
📌 Required Qualifications
7 to 10 years of professional software development experience building scalable, distributed applications or platforms.
Strong experience with Java
Proven experience leading the architecture and delivery of complex, production-critical systems without relying on direct authority.
Deep understanding of distributed systems, concurrency, resiliency, failure handling, data structures, and algorithms.
Experience designing REST, gRPC, GraphQL, and asynchronous service integrations.
Hands-on experience with Kubernetes, containers, and cloud-native architectures.
Experience with observability, incident management, and participating in on-call rotations to resolve complex production issues.
A strong bias for execution while maintaining high standards for quality and reliability.
⭐ Desirable Experience
Experience with AWS, Azure, or Google Cloud Platform.
Experience in incident management, on-call orchestration, SLO/SLI frameworks, incident management practices and AI-assisted root cause analysis.
Experience building internal developer platforms or reliability tooling.
Customer-facing experience representing technical products to enterprise stakeholders.
Strong architectural judgment balancing near-term delivery with long-term scalability, extensibility, and enterprise security.
Experience designing the complete incident lifecycle—from alert ingestion through resolution and post-mortem automation.
Ability to translate complex requirements into scalable AI investigator, event-processing, and integration platforms.
Bachelor’s degree in Computer Science or a related discipline; an advanced degree is preferred.
Equivalent professional experience will also be considered.
🎁 Benefits
Competitive
🛂 Visa & Eligibility
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex or national origin.
Please let Harnessinc know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.