Senior/Staff Software Engineer (Platform and Execution Model)
Red Cell
· Seattle, WA (Preferred) or McLean, VA or Remote (USA)
· Full-time
Compensation$180,000 - $240,000
LocationSeattle, WA (Preferred) or McLean, VA or Remote (USA)
EmploymentFull-time
Role typeEngineering
PostedOct 08, 2026
AWSPlatformCI/CDIaCCloudMLAGIAIObservabilitySRE
📋 Job Overview
Trase Systems is AI, Uncomplicated. Trase empowers enterprise leaders to harness the full potential of AI without the associated complexity and risks. We are an end-to-end solution for deploying, managing, and optimizing AI in the enterprise. Our platform specializes in bridging the “last mile” of AI adoption, unlocking AI's full potential while driving efficiency and significant cost savings.
🏢 About Red Cell Partners
Red Cell Partners is an incubation firm building and investing in rapidly scalable technology-led companies that are bringing revolutionary advancements to market in three distinct practice areas: healthcare, cyber, and national security. United by a shared sense of duty and deep belief in the power of innovation, Red Cell is developing powerful tools and solutions to address our Nation’s most pressing problems.
🎯 The Role
As a Senior or Staff Software Engineer, you'll build and own critical parts of Trase OS, the shared platform that powers Trase deployments in regulated environments. You'll work on the distributed systems and platform primitives connecting workflows, agents, tools, and product surfaces, with a particular focus on reliability, scalability, and correctness.
You'll take on technically ambiguous problems, design solutions, and drive them through production. This is a hands-on engineering role for someone who is comfortable going deep on distributed systems while helping other engineers make sound architectural decisions.
Clean abstractions and correctness-under-failure are critical because we operate long-lived agents in healthcare/defense environments where auditability and reliability are non-negotiable.
The level will reflect your experience and demonstrated scope. Staff-level candidates will be expected to lead architecture across teams, establish engineering standards, and mentor other engineers.
✅ Key Responsibilities
Design, build, and own critical components of the core execution model (state machine, lifecycle, resource model, failure semantics)
Build reliable distributed systems that behave predictably across retries, restarts, partial failures, and concurrent execution.
Guarantee correctness via idempotency, deterministic replays, compensating actions, and data integrity
Engineer reliability at scale: concurrency controls, rate limits, backpressure, sharding/partitioning, and workload isolation
Build security & governance into the core: RBAC/ABAC, policy enforcement, fine-grained audit & lineage
Deliver observability: distributed tracing, structured logs, metrics, and evaluation hooks; build an “explainable trail” of agent actions
Own quality: design reviews, test strategy (unit, property, chaos), performance baselines, SLOs, incident response, and postmortems
Mentor engineers, contribute to design reviews, and help establish strong engineering patterns across the platform.
📌 Required Qualifications
8+ years of experience building distributed/platform systems, including significant experience defining architecture across teams or domains
4+ years owning mission-critical runtimes or workflow/orchestration systems
Deep expertise with durable execution (e.g., state machines, event sourcing, saga/compensation, idempotency, exactly/at-least-once semantics)
Proven track record with security & governance in production systems (auth, RBAC, audit, policy)
Hands-on with observability (Grafana or equivalent), including trace correlation across async boundaries
Strong systems design across storage, queues, schedulers, and evented architectures; performance tuning under load
Excellence in a modern language (e.g., Go, Rust, Java, or TypeScript) and cloud-native stacks (containers, CI/CD, IaC)
Comfortable operating in regulated or high-assurance environments; bias toward correctness, clarity, and documentation
Strong technical judgment and the ability to influence design decisions within a team and across closely related engineering areas.
Ability to incorporate advance LLM capabilities into system design and platform architecture decisions where appropriate
⭐ Desirable Experience
Prior work on workflow engines (Temporal/Cadence/AWS Step Functions, Argo, Airflow) or serverless runtimes
Experience with policy engines (OPA), secrets/KMS, or data-handling controls (PII/PHI)
ML/LLM evaluation frameworks, tool/plugin architectures, or embedding model governance into execution
Government or healthcare experience (HIPAA, audit readiness) and multi-tenant isolation
🎁 Benefits
Career track opportunity with potential for rapid advancement with strong performance as the firm grows
100% employer paid, comprehensive health care including medical, dental, and vision for you and your family.
Paid maternity and paternity for 14 weeks at employees' normal pay.
Unlimited PTO, with management approval.
Opportunities for professional development and continued learning.
Optional 401K, FSA, and equity incentives available.
Mental health benefits are available through Tara Mind.
Cost effective GLP-1 solutions available through Crux.
🛂 Visa & Eligibility
We’re an Equal Opportunity Employer: You’ll receive consideration for employment without regard to race, sex, color, religion, sexual orientation, gender identity, national origin, protected veteran status, or on the basis of disability.
Please let Red Cell know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.