logo

AI Job Summary

Design and lead the architecture for reliability, observability, and health systems on a globally distributed AI infrastructure platform. You'll define technical strategy across runtime, routing, capacity, and specialized hardware, establish production validation practices, and mentor engineers. This role requires six years of software engineering experience with languages like C++, Java, or Python, plus expertise in distributed systems, observability, and load balancing. Work is based in the U.S., with higher pay ranges for the San Francisco Bay Area and New York City. Base salary ranges from $142,800 to $274,800 nationally, or $188,000 to $304,200 in specified metro areas.

Written from this posting by Neural Jobs AI. The full description is below.

AI Resume Tailoring Sign in to use this AI Cover Letter Sign in to use this

Job Description

Overview
We are looking for a Principal Software Engineer to advance the reliability and observability architecture of a large-scale, globally distributed AI platform. You will set technical direction and solve complex challenges across runtime, routing, capacity, infrastructure, and specialized hardware.
 
This role offers the opportunity to work at the intersection of distributed systems, cloud infrastructure, and AI inference. You will learn how advanced AI models are deployed and operated globally, collaborate with experts across the stack, and build foundational capabilities that directly improve customer experience.
 
This is an ideal role for a technical leader who enjoys solving ambiguous, cross-system problems, mentoring engineers, and shaping engineering practices across an organization.
 
Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees, we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day, we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.
 


Responsibilities
Responsibilities
  • Define and drive the architecture and technical roadmap for platform reliability, observability, and operational health.
  • Design unified health and telemetry capabilities that connect customer impact with application, capacity, dependency, infrastructure, and deployment signals.
  • Develop platform safeguards for overload protection, capacity management, routing integrity, configuration consistency, and automated isolation and recovery.
  • Establish engineering practices for production validation, progressive delivery, regression detection, fault testing, and automated rollback.
  • Advance end-to-end request tracing and diagnostics across distributed services, including routing, retries, failover, and asynchronous operations.
  • Lead cross-team architecture efforts, mentor engineers, and turn production learnings into reusable platform capabilities and measurable reliability improvements.


Qualifications

Required Qualifications: 

  • Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
  • OR equivalent experience.

Preferred Qualifications: 

  • Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
    • OR Bachelor's Degree in Computer Science or related technical field 
    • AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
    • OR equivalent experience. 
  • Experience with AI/ML serving platforms, high-performance computing, accelerator-based infrastructure or other compute-intensive distributed systems.
  • Experience with distributed tracing, capacity management, load balancing, admission control, retries, backpressure, and graceful degradation.
  • Proficiency in Azure Monitoring systems in the Azure ecosystem would be a plus.
  • Experience designing, building, and operating large-scale distributed systems, cloud services, or other complex production platforms.
  • Experience with reliability engineering, observability, service health, telemetry, and production incident response.
  • Demonstrated ability to lead complex technical initiatives and drive alignment across engineering teams and organizational boundaries.
  • Strong written and verbal communication skills, including the ability to explain technical strategy and architectural decisions to engineers and senior leaders.
 
#AIINFRA


Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.




Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Do you match this job?

Here is what this employer asked for. Sign in and we will fill in your half.

  • Role Principal Software Engineer - AI Infrastructure
  • Experience 8-9 years
  • Education Bachelor Degree
  • Salary 143K - 275K Yearly
  • Work type Hybrid
  • Location United States
Check my match (free)
Microsoft
Hardware & Semiconductors · 500+ Members · Redmond, WA, United States

Microsoft builds Windows, Azure, Office and the Copilot family of AI assistants, and operates one of the largest AI training and inference fleets in the world. Microsoft Research and the AI platform teams work across foundation models, systems for large-scale training, and applied ML in every product line.

Founded in 1975 and headquartered in Redmond, Washington, the company is also OpenAI's principal compute partner and ships AI tooling for developers through GitHub, VS Code and Azure AI.

39 more Infrastructure & Platform Engineer roles in Mountain View

Job Overview
Salary
143K - 275K Yearly
Eligibility
United States Right to work in the United States required.
Workplace
Hybrid
Job Posted:
3 days ago
Job Type
Full Time
Education
Bachelor Degree
Experience
8-9 years

Share This Job: