Applications are completed on the employer's own site.
About this role
Overview
At Microsoft AI, we are building frontier models and systems that amplify human potential and reach millions of people through Copilot and the broader Microsoft ecosystem. The Reinforcement Learning Engineering team develops the environments, data, evaluations, training systems, and algorithms used to improve our models after pre-training. Our work advances capabilities including reasoning, instruction following, tool use, coding, and agentic knowledge work. We operate across the full reinforcement learning loop: turning product and research goals into measurable tasks, collecting high-quality experience and feedback, running training at scale, evaluating model behavior, and rapidly translating results into the next generation of models.
As a Member of Technical Staff - Reinforcement Learning Engineering, you will combine strong software engineering with applied machine learning. You will set technical direction, lead ambitious research-to production initiatives, and build reliable systems for reinforcement learning at frontier-model scale. You will work closely with researchers, evaluation teams, infrastructure engineers, and product partners to turn ambiguous model and product goals into measurable capability improvements. This is a hands-on role for an engineer who enjoys moving between algorithms and systems, forming clear hypotheses, and owning programs from initial strategy through production impact.
We are looking for candidates who:
- Are passionate about advancing reinforcement learning and post-training for large language models.
- Thrive in a highly collaborative, fast-paced environment with ambitious goals.
- Bring a high degree of craftsmanship and attention to detail.
- Use rigorous experimentation and measurement to make decisions.
- Take end-to-end ownership and are motivated by shipping models and capabilities to users.
Responsibilities
- Set technical direction and lead multi-quarter initiatives to improve large languagemodels and agents through reinforcement learning.
- Build and scale RL environments, rollout systems, training pipelines, and data-generation workflows.
- Develop high-quality tasks, feedback signals, reward models, and training datasets using human, model-generated, and real-world interaction data.
- Form hypotheses and run experiments to improve reasoning, instruction following, tool use, coding, and other agentic capabilities.
- Partner with evaluation teams to define benchmarks, diagnose model failures, and measure both core capabilities and real-world performance.
- Improve the reliability, observability, throughput, and reproducibility of large-scale RL experiments.
- Analyze experiment results, communicate findings, and use them to shape model, data, evaluation, and product roadmaps.
- Collaborate closely with research, infrastructure, safety, and product teams to translate new techniques into models used by millions of people.
- Mentor engineers, raise the quality of experimental and system design, and contribute to a positive, inclusive engineering culture grounded in knowledge sharing and technical excellence.
Required/Minimum Qualifications
- Bachelor's Degree in Computer Science or a related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python, OR equivalent experience.
Additional/Preferred Qualifications
- Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
- Experience building reliable machine learning systems and applying reinforcement learning, reward modeling, preference optimization, model post-training, or closely related techniques to large language models or other large-scale generative models.
- Demonstrated experience leading technically ambitious, multi-team machine learning initiatives from initial hypothesis and experimental design through measurable model or product impact.
- Experience influencing technical direction across organizational boundaries and mentoring engineers, including using quantitative evidence to make decisions and deliver high-quality results through ambiguity.
- Experience building RL environments, agent harnesses, simulators, or large-scale rollout infrastructure.
- Experience with distributed training or inference systems using accelerators such as GPUs.
- Experience generating, curating, filtering, and versioning large training datasets.
- Experience developing reward signals from human feedback, model feedback, verifiers, or real-world outcomes.
- Experience evaluating agents on reasoning, coding, tool use, or long-horizon tasks.
- Familiarity with modern post-training methods and open-source machine learning frameworks.
- Demonstrated written and verbal communication skills, including explaining complex experimental results to cross-functional partners.
- Experience shaping an RL, post-training, or evaluation roadmap across multiple teams or model releases.
Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.