logo

AI Job Summary

This role involves developing LLM evaluation strategies, measurement frameworks, and responsible AI implementations for Outlook's Copilot features. You'll need a bachelor's degree in a relevant field plus 4+ years in statistics, predictive analytics, or research (or equivalent with advanced degrees). Key skills include data science, machine learning, experimentation, LLM evaluation techniques, and Python. The position is based in the U.S., with base pay ranging from $119,800–$234,700 annually, or $160,200–$261,000 in the San Francisco Bay Area and New York City.

Written from this posting by Neural Jobs AI. The full description is below.

AI Resume Tailoring Sign in to use this AI Cover Letter Sign in to use this

Job Description

Overview
The Outlook team at Microsoft is reimagining how people communicate, organize their work, and manage time through intelligent, trustworthy experiences across Mail and Calendar. We build products used at global scale, including Copilot-powered experiences for drafting, summarization, search, inbox prioritization, meeting preparation, scheduling, and agentic workflows. Our work spans product analytics, experimentation, telemetry, machine learning, generative AI, and large language model evaluation to improve customer value, quality, trust, and adoption across Outlook experiences. 
 
As a Senior Applied Scientist, you will play a critical role in advancing our Outlook's Copilot efforts in the areas of Large Language Model (LLM), Evaluation, Relevance and Responsible AI (RAI). This multifaceted role is responsible for developing an end-to-end infrastructure and measurement framework, fostering cross-functional collaboration, and leveraging data science and AI experience to guide decision-making. The successful candidate will work with multiple large organizations and stakeholders to drive the evaluation of our LLM systems and associated components. 
 
Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.


Responsibilities
  • Strategic Leadership: Develop and execute a comprehensive strategy for LLM evaluation, encompassing LLM quality, costs, model performance, model utility (user experience and prompt effectiveness), and responsible AI considerations, in alignment with company-wide efforts and informed by emerging research. Propose and drive cost effective solutions such as small language models (SLMs), and finetuned models and ensemble models. 

  • Program Management: Oversee and manage large-scale, cross-functional evaluation programs, ensuring alignment with organizational objectives and timelines. Develop and maintain a robust measurement framework to track and report on LLM performance and user impact. Drive engineering product roadmap to construct automated evaluation pipelines integrated into the product workflow. 

  • Applied Science Experience: Utilize data science skills to design experimentation, analyze data, create OKRs, create measurement and metrics, and derive actionable insights to enhance LLM systems. Responsible for influencing product and user experience based on evaluation results.  

  • Model and Agent Evaluation: Lead efforts to assess and improve the performance and effectiveness of language models and agents, driving iterative enhancements, including synthetic and manufactured data creation. 

  • User Experience Enhancement: Collaborate with User Experience teams to evaluate and optimize user interactions with AI systems, enhancing user satisfaction. 

  • Responsible AI (RAI): Implement RAI and DSB principles and guidelines in AI systems, ensuring ethical and unbiased practices in model development and deployment. 

  • Contribute to the LLM research body: Form partnerships and lead deep research initiatives in areas of LLM evaluation and user experience optimization that contribute to the scientific body and deepen the product team's understanding and expertise of user mental models of and alignment to LLM-powered experience.

  • Cross-Functional Collaboration: Work with engineering, research, product, and other teams to ensure seamless integration of evaluation processes into the development lifecycle. 

  • Stakeholder Engagement: Communicate findings and recommendations to executive leadership, fostering a data-driven culture within the organization. 



Qualifications

Required Qualifications:

  • Bachelor's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 4+ years related experience (e.g., statistics predictive analytics, research)
    • OR Master's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 3+ years related experience (e.g., statistics, predictive analytics, research)
    • OR Doctorate in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 1+ year(s) related experience (e.g., statistics, predictive analytics, research)
    • OR equivalent experience. 

Other Requirements:
Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings:
 
Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.

Preferred Qualifications:

  • Ph.D. or Master’s in a relevant field (e.g., Data Science, Computer Science, Social science, HCI, etc.) 
  • Experience in data science, machine learning, experimentation, and AI, with a track record of delivering impactful results. 
  • Experience in program management and leading cross-functional teams. 
  • Experience with LLM in finetuning, reinforcement learning, evaluation techniques, implementing RAG techniques, agentic workflows and industry best practices. 
  • Development experience in Python 
  • Analytical, problem-solving and presentation skills. 
  • Understanding of responsible AI principles. 
  • Ability to work in a fast-paced and dynamic environment. 

 

#Outlook; #OPG 



Applied Sciences IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.




Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Do you match this job?

Here is what this employer asked for. Sign in and we will fill in your half.

  • Role Senior Applied Scientist - Outlook Science Team
  • Experience 5-7 years
  • Education Bachelor Degree
  • Salary 120K - 235K Yearly
  • Work type Hybrid
  • Location United States
Check my match (free)
Microsoft
Hardware & Semiconductors · 500+ Members · Redmond, WA, United States

Microsoft builds Windows, Azure, Office and the Copilot family of AI assistants, and operates one of the largest AI training and inference fleets in the world. Microsoft Research and the AI platform teams work across foundation models, systems for large-scale training, and applied ML in every product line.

Founded in 1975 and headquartered in Redmond, Washington, the company is also OpenAI's principal compute partner and ships AI tooling for developers through GitHub, VS Code and Azure AI.

All jobs at Microsoft
Job Overview
Salary
120K - 235K Yearly
Eligibility
United States Right to work in the United States required.
Workplace
Hybrid
Job Posted:
5 days ago
Job Type
Full Time
Education
Bachelor Degree
Experience
5-7 years

Share This Job: