Research Services👥 201 employees📍 San Francisco, CA, USEst. 2015
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. AI is an extremely powerful tool that must be created with…
📋 Job Overview
Own the end-to-end data-center hardware quality and reliability loop for OpenAI’s 3P infrastructure and 1P current and next-gen platforms. Turn field failures into quantified risk, fast containment, verified root cause, improved MQE/NPI and manufacturing-test coverage, accurate spares forecasts, and upstream changes that prevent recurrence.
🏢 About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the
🎯 The Role
The first hire must combine practical hardware/system understanding, reliability engineering, data fluency, and cross-functional technical leadership at data-center scale.
✅ Key Responsibilities
Build and govern the field-quality data model across telemetry, tickets, RMA/repair, FA, firmware, configuration, supplier, and manufacturing genealogy.
Define AFR, ASR, DPPM, MTBF/MTTR, repeat-repair, NTF, repair-cycle-time, and forecast-versus-actual metrics with explicit denominators and uncertainty.
Provide fleet-level macro views and unit/FRU/cohort-level micro views; detect shifts and bound affected populations.
Lead systemic field-failure triage, containment, failure analysis, 8D/CAPA, risk assessment, corrective-action verification, and recurrence monitoring.
Develop cohort, life-data, reliability-growth, and spare-demand projections by product, FRU, supplier, configuration, geography, and age.
Partner with MQE and NPI to convert field mechanisms into manufacturing-test coverage, screening/stress profiles, diagnostics, control plans, DFR/DFS requirements, FMEA/FTA, mission profiles, FRU strategy, and qualification gates.
Close the loop by verifying whether upstream changes reduce field recurrence.
Define supplier/CM FA standards, field-data contracts, scorecards, escalation paths, and closure evidence.
Provide serviceability, TCO, and spares inputs without owning inventory execution or procurement.
Create concise executive decision packages: population at risk, exposure, confidence, options, cost/risk, and recommendation.
Run the cross-functional reliability council and, as the team grows, mentor the 1P and 3P Field Quality Engineers.
📌 Required Qualifications
BS in electrical, mechanical, computer, materials, reliability engineering, physics, or equivalent experience; MS preferred.
8+ years in hardware quality/reliability, server/rack systems, or mission-critical infrastructure; 3+ years owning field-failure, RMA, or CAPA outcomes.
Solid working understanding of hardware and system architecture across board, tray, rack, firmware, telemetry, manufacturing test, and fleet behavior; deep expertise in every subsystem is not required.
Reliability statistics: censored life data, Weibull/Poisson/binomial methods, confidence bounds, MTBF/MTTR, and reliability growth.
Hands-on FMEA/FTA, accelerated or reliability-demonstration testing, 8D/CAPA, FA, and corrective-action verification.
Working proficiency with SQL and Python/R or equivalent analytics tools.
Ability to influence design, validation, operations, suppliers/CMs, and senior leaders without direct authority.
⭐ Desirable Experience
GPU/AI server platforms, liquid cooling, high-power delivery, high-speed networking, rack integration, or data-center operations.
Design for serviceability: FRU boundaries, diagnostics, repair workflows, tooling/access, and spares policy.
Qualification-to-field correlation and mission-profile development.
ODM/CM/supplier experience: FA quality, audit, QBR, and corrective-action governance.
Linux/BMC/IPMI/Redfish logs and fleet telemetry.
Leadership of a cross-generation reliability program or launch-readiness gate.
🎁 Benefits
Compensation: $226K – $285K • Offers Equity
Medical, dental, and vision insurance for you and your family, with employer contributions to Health Savings Accounts Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses (parking and transit) 401(k) retirement plan with employer match Paid parental leave (up to 24 weeks for birth parents and 20 weeks for non-birthing parents), plus paid medical and caregiver leave (up to 8 weeks) Paid time off: flexible PTO for exempt employees and up to 15 days annually for non-exempt employees 13+ paid company holidays, and multiple paid coordinated company office closures throughout the year for focus and recharge, plus paid sick or safe time (1 hour per 30 hours worked, or more, as required by applicable state or local law) Mental health and wellness support Employer-paid basic life and disability coverage Annual learning and development stipend to fuel your professional growth Daily meals in our offices, and meal delivery credits as eligible Relocation support for eligible employees Additional taxable fringe benefits, such as charitable donation matching and wellness stipends, may also be provided.
🛂 Visa & Eligibility
This role is at-will and OpenAI reserves the right to modify base pay and other compensation components at any time based on individual performance, team or company results, or market conditions.