Applications are completed on the employer's own site.
Summary
Design and operate Kubernetes-based distributed systems to manage GPU compute infrastructure across data centers.
About this role
Software Engineer, Compute Foundations
San Francisco, California, United States·Posted Today
full-time
LinkedIn
Apply now
The Role
OpenAI's Compute Foundations team builds software that manages GPU compute infrastructure across data centers and sites, supporting model training and inference. In this role, you will design and operate Kubernetes-based distributed systems that provision, configure, and manage compute resources throughout their lifecycle, connecting global services with bare-metal systems management.
What You'll Do
Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites and scale as GPU capacity grows
Define APIs and resource models that enable clients to request and track lifecycle operations across diverse hardware platforms and providers
Build provisioning and configuration services that coordinate network boot, hardware management interfaces, firmware deployment, operating-system images, drivers, and host configuration
Develop lifecycle management systems for discovery, allocation, provisioning, upgrades, maintenance, recovery, and decommissioning, integrated with health and validation systems
Design reliable reconciliation and recovery mechanisms for concurrent changes, interrupted operations, and partial failures with staged rollouts across nodes, racks, and clusters
What You Need
Strong software engineering fundamentals with experience designing, implementing, and owning production distributed systems or infrastructure services
Experience developing infrastructure systems that use Kubernetes APIs and reconciliation to manage resources
Understanding of bare-metal node provisioning from power-on to configured workload-ready state, with depth in areas such as PXE, DHCP/DNS, baseboard management controllers (BMCs), firmware, Linux, drivers, images, or configuration management
Ability to design reliable APIs and asynchronous workflows, reasoning about concurrency, consistency, idempotency, and failures across service and provider boundaries
Capability to diagnose reliability and performance problems across service, operating-system, and machine boundaries
Nice to Have
Experience building infrastructure control planes that coordinate operations across multiple sites or regions
Prior work with GPU or HPC infrastructure, including topology and shared dependencies across machines, racks, or clusters
Integration experience with multiple hardware platforms or infrastructure providers into a common service or resource model
Skills named in this posting
Kubernetes
Distributed Systems
Linux
Bare-metal Provisioning
API Design
Firmware Deployment
Hardware Management
About Linuxconfig
Linuxconfig provides tutorials and guides for Linux users.
Automation Machinery Manufacturing
Similar Software Engineer, Compute Foundations jobs in San Francisco, CA