Fluidstack is seeking a Reliability Engineer, Data Center Design to join our team. We exist to make humanity more free by delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators.
Extreme ownership. Full autonomy. Own things end to end often taking on scope outside your core role without being asked to get things done.
Velocity. We drive everything forward as fast as possible.
First principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.
Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.
Examples of key problems the team is working on:
Build and maintain reliability block diagrams and fault tree analyses across the full MEP service chain, from utility intake to rack-level IT load, running Monte Carlo simulations of at least 100,000 iterations to produce P50/P90/P95/P99 availability distributions.
Own the reliability study at each 30% and 90% design gate for Fluidstack's own templates, and run independent reliability assessments of EPC-proposed, colocation, and acquired-site designs benchmarked against Uptime Institute Tier III/IV classifications.
Manage third-party reliability consultants and review their RBD, FTA, and Monte Carlo work for methodology soundness, turning identified single points of failure into prioritized design recommendations before capital is committed.
Maintain the reliability model, component data, and full audit trail inside Windchill PLM, and produce availability statements, sensitivity analyses, and FMEA summaries for lease documents, SLAs, and investor materials.
Feed live-site failure and repair data from Operations and Commissioning back into the models, track modeled availability against measured uptime across the portfolio, and translate the components driving unavailability into maintenance, sparing, and capital allocation recommendations such as N+1 versus 2N.
You've personally built reliability block diagrams and fault tree analyses for mission-critical electrical and mechanical systems, not modeled them in the abstract.
You've run Monte Carlo simulations for system availability using PTC Windchill Prediction, ReliaSoft, or equivalent tools, and you know IEEE 493 (Gold Book) and IEEE 3006.5 well enough to defend your failure-rate and repair-time assumptions.
You've managed reliability models, component data, and version control inside PTC Windchill as the system of record, not a spreadsheet on the side.
You understand data center MEP systems well enough to model them accurately: MV/LV electrical distribution, standby generation, UPS, chilled water plants, CDUs, and building management/controls systems.
You default to quantifying risk instead of describing it. You'd rather hand someone a P90 availability number than tell them a system should be reliable.
You translate technical reliability findings into figures a lease document, SLA, or investor deck can actually use, without losing what the number means.
Bonus: PE license. Direct experience benchmarking designs against Uptime Institute Tier III/IV classifications. Liquid or hybrid cooling reliability modeling. Managing outside reliability consultants or engineering firms on a deliverable basis.
Join 2,000+ DevOps developers getting weekly alerts for remote and US/EU roles, Kubernetes, AWS, Terraform, filtered for your stack.
Get a dedicated static EU outbound IP for Banks, payments, EHRs, APIs, AI.
Get instant access to exclusive DevOps jobs with €120K+ salaries
Best value for job search
Only €4.13/month - Save 75%