logo
AI Resume Tailoring Sign in to use this AI Cover Letter Sign in to use this

Job Description

Overview
The AI Infrastructure team is responsible for building and operating the large-scale, reliable, and efficient GPU-based clustering infrastructure that powers Microsoft’s AI/ML ecosystem. We host the training and inference platforms behind many of Microsoft’s flagship AI offerings, including Azure OpenAI Service, M365 Copilot & Copilot Tuning, GitHub Copilot, Azure AI Foundry’s inference and fine-tuning services for both OpenAI and open-source models; as well as the mature Azure ML Services, which provide data scientists and developers a rich experience for defining, training, fine-tuning, deploying, monitoring, and consuming machine learning models. Our infrastructure enables AI innovation at hyperscale and supports some of the most demanding workloads and business groups across the company. We engage directly with some of the major internal research and applied AI/ML groups using these services, including Microsoft Research, M365, Microsoft Security, and the Bing WebXT team.
 
The AI Infra team is looking for a talented Software Engineer II, with initial focus on the Scheduler subsystem. The scheduler is the “brains” of the AI Infra control plane. It governs access to the GPU, NPU and CPU capacity of the platform according to a complex system of workload preference rules, placement constraints, optimization objectives, and dynamically interacting policies aimed to maximize hardware utilization and fulfill greatly varying needs of users and the AI platform partner services in terms of workload types, prioritization, and capacity targeting flexibility. The scheduler’s set of capabilities is broad and ambitions. It manages quota, capacity reservations, SLA tiers, preemption, auto-scaling, and a wide range of configurable policies. It is both a workload-aware and topology-aware scheduler (down to the level of cluster racks and nodes). Global scheduling is a distinctive major feature that overcomes the regional segmentation of the Azure compute fleet by treating the GPU capacity as a single global virtual pool, which greatly increases capacity availability and utilization for major classes of AI/ML workload. We have achieved this capability by avoiding a significant global single point of failure, based on regional instances of the scheduler service interacting via peer-to-peer protocols for sharing capacity inventory and coordinating handoff of jobs for scheduling. Our system manages significant amount of GPU capacity even outside Azure datacenters, through a unified model and operational process and highly generalized, flexible workload scheduling capabilities.
 
To be able to manage the inherent complexity of the Scheduler subsystem and enable it to meet the stringent expectations of high service reliability, availability, and throughput, we emphasize rigorous engineering, utmost precision and quality, and strong ownership—from feature design to livesite. Quality mindset, attention to detail, development process rigor, and data-driven design and problem-solving skills are key for success in our mission-critical control plane space. We enjoy great creative freedom and thrive on the capacity management and workload scheduling challenges posed to us by the various partner teams.
 

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.



Responsibilities
  • Work on the design and development of the core AI Infrastructure distributed and in-cluster services that support large scale AI training and inferencing.
  • Develop, test, and maintain control plane services written in C#, hosted on Service Fabric or Kubernetes (AKS) clusters.
  • Enhance systems and applications to ensure high stability, efficiency and maintainability, low latency, tight cloud security.
  • Provide operational support and DRI (on-call) responsibilities for the service.
  • Develop and foster a deep understanding of the AI/ML concepts, use cases, and relevant services used by our customers. Be an AI-first developer, making productive use of the available tools and actively engaging in experimentation and learning.
  • Collaborate closely with service engineers, product managers, and internal applied research and data science teams within Microsoft to build better solutions together.
  • Provide vision, expertise, and technical leadership to other team members.
  • Embody our culture and values
 


Qualifications

Required Qualifications: 

  • Bachelor's Degree in Computer Science or related technical field AND 2+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
    • OR equivalent experience. 

Other Requirements:

Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:

Microsoft Cloud Background Check:

- This position will be required to pass the Microsoft background and Microsoft Cloud background check upon hire/transfer and every two years thereafter.

Preferred Qualifications: 

  • Hands-on (devops) experience with larger-scale, high-availability cloud services at the PaaS or IaaS level, based on microservices architecture, ideally related to AI infrastructure or workload hosting
  • Proficiency with use of complex data structures and algorithms, preferably in the setting of a resource allocator/scheduler, workflow/execution orchestration engine, database engine, or similar
  • Proficiency and thoroughness in unit testing and testability techniques
  • Agentic development skills
  • Experience with building and operating “stateful” and critical control plane services; handling challenges with data size and data partitioning; advanced use of a NoSQL cloud database
  • Service reliability and fundamentals engineering; instrumentation for KPIs or performance analysis; demonstrated service and code quality mindset
  • Applied knowledge of Kubernetes: service model, workload packaging and deployment, programmatic extensibility (CRDs, operators); or equivalent knowledge of Service Fabric; experience with any service mesh
  • Data-driven design and troubleshooting and data analytics skills, ideally with Kusto
#AIINFRA
 


Software Engineering IC3 - The typical base pay range for this role across the U.S. is USD $102,100 - $202,200 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $133,800 - $219,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.




Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Do you match this job?

Here is what this employer asked for. Sign in and we will fill in your half.

  • Role Software Engineer II
  • Experience 3-4 years
  • Education Any
  • Salary 102K - 202K Yearly
  • Work type Hybrid
  • Location United States
Check my match (free)
Microsoft
Hardware & Semiconductors · 500+ Members · Redmond, WA, United States

Microsoft builds Windows, Azure, Office and the Copilot family of AI assistants, and operates one of the largest AI training and inference fleets in the world. Microsoft Research and the AI platform teams work across foundation models, systems for large-scale training, and applied ML in every product line.

Founded in 1975 and headquartered in Redmond, Washington, the company is also OpenAI's principal compute partner and ships AI tooling for developers through GitHub, VS Code and Azure AI.

All jobs at Microsoft
Job Overview
Salary
102K - 202K Yearly
Eligibility
United States Right to work in the United States required.
Workplace
Hybrid
Job Posted:
1 day ago
Job Type
Full Time
Education
Any
Experience
3-4 years

Share This Job: