At SF Tensor, we're building the future of high-performance compute. We firmly believe that the future of AI depends on rethinking and rebuilding the stack, from hardware to compiler to cloud. We're here to make compute faster, cheaper, and more available. We're building our Kernel Optimizer and Model Foundry to achieve this. We're backed by Susa Ventures, Y Combinator, and other great funds and angels.
🏢 About SF Tensor
We build the fastest GPU compiler in the world. Our approach allows us to search a far wider space while still guaranteeing correctness. We hold #1 on NVIDIA's own kernel benchmark across hundreds of production kernels. We're looking for researchers, engineers, and organizations who agree with the basic premise: you don't get the next leap in AI without a leap in compute first.
🎯 The Role
We're hiring a Member of Technical Staff for Sandbox Infrastructure to build a serverless GPU container service across NVIDIA, AMD, TPU, and Trainium. This service needs to run our compiler, post-training runs, and customer workloads that require isolation. The work now is depth and breadth, adding support for more vendors, higher fidelity instrumentation, faster cold starts, and more features.
✅ Key Responsibilities
Extend our sandboxing stack to new vendors and accelerators
Build and maintain GPU virtualization below the runtime, including gVisor work at the driver and ioctl level
Make sandboxes first-class citizens on spot capacity, which means preemption-aware scheduling, checkpointing, and rescheduling
Support multi-GPU and multi-node sandboxes, including the interconnect (NVLink, NVSwitch) and RDMA paths (InfiniBand, RoCE) paths those require
Own live migration end-to-end, including our socket-preserving migration
Guarantee measurement and profiling fidelity as well as their isolation
Work directly with the compiler, post-training, and kernel teams to ensure their throughput is not capped by sandboxes
📌 Required Qualifications
Strong low-level systems engineering background: Linux kernel internals, containers, namespaces, cgroups, syscall interception, or hypervisors
Experience in GPU systems engineering: drivers, runtimes, or scheduling on accelerator fleets
Comfortable with distributed systems failure modes: preemption, partial failure, checkpoint/restore, and dealing with states you can't afford to lose
Proficient in Go, C/C++, or Rust
Strong bias toward building the thing yourself when no vendor supports what you need
⭐ Desirable Experience
Worked directly with gVisor, Firecracker, Kata, QEMU/KVM, or similar
Worked directly on CRIU, live migration, or connection-preserving failover work
Familiar with NCCL/RCCL, RDMA, InfiniBand, or vendor interconnects
Run large fleets on spot or other preemptible capacity
Familiar with bare-metal provisioning, hypervisors, or fleet management at scale
Security background in isolation boundaries and untrusted code execution
🎁 Benefits
Salary: $275K – $315K • Offers Equity
🛂 Visa & Eligibility
We believe that hard problems get solved in person and most of our work happens at our office in San Francisco.
Please let Sf Tensor know that you found this role at devopsprojectshq.com as a way to support us, so we can keep providing you with awesome DevOps jobs.
Never miss a job
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
🔒 Need an IP to whitelist?
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.