logo
AI Resume Tailoring Sign in to use this AI Cover Letter Sign in to use this

Job Description

  • Architect the GPU Reliability Systems Operations and Tooling SW organization to transition from NPI-specific task forces into a centralized organization.
  • Drive the GPU support strategy by developing a roadmap for fleet health reliability, capacity turn-up, and automated health management.
  • Establish and enforce Service Level Objectives and Service Level Indicators for GPU Pod availability and performance.
  • Sponsor automation and toil reduction efforts by driving the development of advanced telemetry and debugging tooling.
  • Manage high-severity escalations for critical hardware and software issues, leading incident response and blameless post-mortems.

Minimum qualifications:

  • Bachelor’s degree, or equivalent practical experience.
  • 8 years of experience programming in C++, Java, Python, Kotlin or Go.
  • 5 years of experience with software architecture and embedded systems.
  • 3 years of experience in a technical leadership role.
  • 2 years of experience in a people management or team leadership role.
  • Experience with GPU programming, systems reliability and computer architecture.

Preferred qualifications:

  • Master's degree or PhD in Computer Science or related technical field.
  • 3 years of experience working in a complex, matrixed organization.
  • Experience in organizational design to consolidate task forces, steering platforms from New Product Introduction to General Availability.
  • Experience operating in a Cloud environment and partnering with strategic customers to improve infrastructure.
  • Experience leveraging AI platforms, Generative AI Agents, and distributed systems for software development.
  • Strong background in GPU Systems Operations, designing automated diagnostic frameworks and telemetry for hardware and software faults.
  • Bachelor’s degree, or equivalent practical experience.
  • 8 years of experience programming in C++, Java, Python, Kotlin or Go.
  • 5 years of experience with software architecture and embedded systems.
  • 3 years of experience in a technical leadership role.
  • 2 years of experience in a people management or team leadership role.
  • Experience with GPU programming, systems reliability and computer architecture.
Do you match this job?

Here is what this employer asked for. Sign in and we will fill in your half.

  • Role Engineering Manager, GPU Reliability, Accelerators
  • Experience 5-7 years
  • Education Bachelor Degree, or equivalent experience
  • Work type On-site
  • Location United States
Check my match (free)
Google
AI Research Lab · 500+ Members · Mountain View, CA, United States

Google builds internet, software, cloud, and AI products used by consumers, developers, and organizations. Its portfolio includes Search, YouTube, Android, Chrome, Maps, Gmail, Workspace, Google Cloud, advertising platforms, devices, and Gemini AI products. The company develops large-scale computing infrastructure and research that power information retrieval, communication, productivity, media, navigation, and machine learning. Google is the largest operating business within Alphabet and earns a substantial share of its revenue from digital advertising.

All jobs at Google

29 more Engineering Manager roles in Sunnyvale

Job Overview

Approx. salary range

177K – 228K

Our estimate — this employer did not publish a salary

Our estimate, not the employer’s. Worked out from the middle half of 215 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.

Eligibility
United States Right to work in the United States required.
Workplace
On-site
Job Posted:
20 hours ago
Job Type
Full Time
Education
Bachelor Degree, or equivalent experience
Experience
5-7 years

Share This Job: