Applications are completed on the employer's own site.
About this role
Overview
At Microsoft AI, we are building frontier models and systems that amplify human potential and reach millions of people through Copilot and the broader Microsoft ecosystem.
The AI Infrastructure and Model Foundry team builds the systems that turn advances in model research into reliable, efficient, and widely usable capabilities. We provide the foundation for the model lifecycle, including experimentation, training and post-training workflows, evaluation, model packaging, deployment, inference, observability, and continuous improvement. Our goal is to let researchers and product engineers move quickly while meeting an exceptional bar for scale, performance, reliability, safety, and operational excellence.
As a Member of Technical Staff - AI Infrastructure and Model Foundry, you will set technical direction and build high-performance distributed systems for frontier models. Depending on your background, you may work across model serving, inference optimization, accelerator orchestration, experiment platforms, developer tooling, or the control and data planes that connect the end-to-end model lifecycle. You will lead ambiguous, high-impact initiatives across organizational boundaries and collaborate closely with model researchers, reinforcement learning engineers, product teams, and hardware and cloud partners to remove bottlenecks and make new model capabilities available at scale.
We are looking for candidates who:
- Are passionate about building the infrastructure behind frontier AI.
- Thrive in a highly collaborative, fast-paced environment with ambitious goals.
- Bring a high degree of craftsmanship and attention to detail.
- Enjoy diagnosing hard systems problems across software, models, and hardware.
- Take end-to-end ownership and are motivated by delivering dependable platforms used by both researchers and customers.
- Balance model and product velocity with disciplined decisions about reliability, latency, throughput, and infrastructure cost.
Responsibilities
- Set technical direction and lead multi-quarter initiatives for scalable frontier-model infrastructure, from experimentation and training through evaluation, deployment, and inference.
- Develop high-throughput, low-latency model-serving systems that efficiently use GPUs and other accelerators.
- Build reusable platform capabilities for model onboarding, versioning, packaging, rollout, routing, autoscaling, monitoring, and rollback.
- Improve inference performance and fleet efficiency through techniques such as batching, prefix and KV caching, workload placement, scheduling, parallelism, quantization, compilation, and memory optimization.
- Create reliable interfaces and developer workflows that help research and product teams move models from experiments to production quickly and safely.
- Define service-level objectives and build observability, benchmarking, capacity-management, and regression-detection systems for model quality, latency, throughput, reliability, and cost.
- Diagnose and resolve bottlenecks across distributed systems, networking, storage, runtimes, model architectures, and accelerator hardware.
- Lead the response to complex production incidents, drive root-cause analysis, and turn operational lessons into durable platform improvements.
- Partner with researchers and reinforcement learning engineers to support large-scale experiment, rollout, evaluation, and data-generation workloads.
- Establish technical and operational guardrails that allow teams to move quickly without creating regressions in availability, security, privacy, performance, or fleet economics.
- Mentor engineers, raise the quality of architecture and design reviews, and contribute to a positive, inclusive engineering culture grounded in knowledge sharing and technical excellence.
Required/Minimum Qualifications
- Bachelor's Degree in Computer Science or a related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Python, Rust, or Go, OR equivalent experience.
Additional/Preferred Qualifications
- Experience designing, building, and operating distributed systems or high-scale production infrastructure, including machine learning infrastructure, model serving, distributed compute, cloud platforms, or another performance-critical systems domain.
- Demonstrated experience leading technically ambitious, multi-team infrastructure initiatives from architecture through dependable production operation and measurable impact.
- Experience influencing technical direction across organizational boundaries and mentoring engineers, supported by strong systems fundamentals in areas such as concurrency, networking, storage, resource management, reliability, or performance analysis.
- Experience serving large language models or other large generative models in production.
- Experience with containers, schedulers, orchestration systems, and large-scale cloud infrastructure.
- Experience building training, post-training, evaluation, or experiment-management platforms.
- Familiarity with model-serving techniques such as continuous batching, prefix or KV caching, disaggregated prefill and decode, speculative decoding, parallelism, quantization, and model compilation.
- Experience operating highly available, GPU-backed services with demanding latency, throughput, capacity, and cost requirements.
- Experience making hardware-software co-design decisions and improving accelerator utilization or fleet economics at substantial scale.
- Demonstrated written and verbal communication skills, including communicating system tradeoffs and operational risks to cross-functional partners.
Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.