logo
AI Resume Tailoring Sign in to use this AI AI Cover Letter Sign in to use this AI

Job Description

Overview

The AI Infrastructure team is responsible for building and operating large-scale, highly reliable, and efficient GPU infrastructure that powers Microsoft’s AI ecosystem. We host the training and inference platforms behind many of Microsoft’s flagship AI offerings, including Microsoft 365 Copilot, GitHub Copilot, Microsoft Copilot, and Azure AI Foundry’s inference and fine-tuning services for both OpenAI and open-source models. Our infrastructure enables AI innovation at hyperscale and supports some of the most demanding workloads across the company.

As a Software Engineer on the AI infrastructure team, you will work on cutting edge infrastructure and tools to support large scale model deployments, pre-training, post-training and fine-tuning on latest generation of NVIDIA and AMD GPUs in Azure and Microsoft partner clouds on some of the world’s largest AI Supercomputers.  

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.



Responsibilities

As an engineer on the AI infrastructure team, your responsibilities include:

•    Design, develop, and maintain AI infrastructure services in Go, Rust, Python, C++, and C#, deployed on large-scale Kubernetes clusters to support inference, pre-training, and post-training workloads for state-of-the-art AI models.
•    Collaborate with engineers, researchers, and external partners to troubleshoot issues, improve reliability, and optimize the performance of large-scale AI training and inference systems.
•    Build and enhance distributed systems that deliver high reliability, low latency, operational efficiency, and strong security across Azure and partner cloud environments.
•    Develop automation and tooling to improve GPU capacity utilization, streamline fleet operations, and enable efficient scaling of AI infrastructure.
•    Provide operational support, technical leadership, and vision while contributing to the deployment, monitoring, and continuous improvement of engineering systems and practices.



Qualifications

Required Qualifications:

  • Bachelor's Degree in Computer Science or related technical field AND 2+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.

 

Preferred Qualifications: 

•    2+ years designing, developing, and shipping high quality software.  
•    2+ years of experience with distributed systems and cloud-based infrastructure. 
•    1+ year of experience with DevOps practices (CI/CD, automated testing, deployment, etc.).   
•    2+ years of software development experience in C#, C++, Python, or similar languages.  
•    2+ years of experience with containerization tools (e.g., Docker, Kubernetes).
•    Knowledge and hands on experience with production ML systems, large-scale training infrastructure, NCCL, CUDA libraries and tools

 

#AIINFRA



Software Engineering IC3 - The typical base pay range for this role across the U.S. is USD $102,100 - $202,200 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $133,800 - $219,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.




Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Do you match this job?

Here is what this employer asked for. Sign in and we will fill in your half.

  • Role Software Engineer II you: —
  • Experience 3-4 years you: —
  • Education Any you: —
  • Salary 102K - 202K Yearly you: —
  • Work type Hybrid you: —
  • Location United States you: —
Check my match
Microsoft
Hardware & Semiconductors · 500+ Members · Redmond, WA, United States

Microsoft builds Windows, Azure, Office and the Copilot family of AI assistants, and operates one of the largest AI training and inference fleets in the world. Microsoft Research and the AI platform teams work across foundation models, systems for large-scale training, and applied ML in every product line.

Founded in 1975 and headquartered in Redmond, Washington, the company is also OpenAI's principal compute partner and ships AI tooling for developers through GitHub, VS Code and Azure AI.

All jobs at Microsoft
Job Overview
Salary
102K - 202K Yearly
Eligibility
United States Right to work in the United States required.
Workplace
Hybrid
Job Posted:
6 days ago
Job Type
Full Time
Education
Any
Experience
3-4 years

Share This Job: