logo
AI Resume Tailoring Sign in to use this AI Cover Letter Sign in to use this

Job Description

  • Design and develop highly scalable software and firmware systems that detect, diagnose, and mitigate reliability issues across Google's server fleet
  • Act as the technical domain expert for the hardware/software boundary, translating complex physical fault behaviors into robust software telemetry, diagnostic, and automated recovery mechanisms.
  • Drive the technical goal for fault management. Guide the engineering team through complex problem-solving and system design, and set the standard for high-quality, reliable code.
  • Partner closely with Hardware Engineering, Technical Infrastructure, and Cloud teams to influence the architectural design of next-generation compute and storage systems for maximum reliability.
  • Architect pipelines to collect and analyze fleet-wide telemetry, turning raw hardware health signals into actionable insights and automated mitigation strategies.

Minimum qualifications:

  • Bachelor's degree or equivalent practical experience.
  • 8 years of experience programming in C++, SQL and embedded systems.
  • 5 years of experience testing, and launching software products.
  • 5 years of experience building and developing large-scale infrastructure, distributed systems or networks, or experience with compute technologies, storage, or hardware architecture.
  • 3 years of experience with software design and architecture.

Preferred qualifications:

  • Master’s degree or PhD in Engineering, Computer Science, or a related technical field.
  • 8 years of experience with data structures and algorithms.
  • 3 years of experience in a technical leadership role leading project teams and setting technical direction.
  • 3 years of experience working in a complex, matrixed organization involving cross-functional, or cross-business projects.
  • Experience with business intelligence platform, SQL pipelines.
  • Bachelor's degree or equivalent practical experience.
  • 8 years of experience programming in C++, SQL and embedded systems.
  • 5 years of experience testing, and launching software products.
  • 5 years of experience building and developing large-scale infrastructure, distributed systems or networks, or experience with compute technologies, storage, or hardware architecture.
  • 3 years of experience with software design and architecture.
Do you match this job?

Here is what this employer asked for. Sign in and we will fill in your half.

  • Role Software Engineer, Google Cloud Platform, Fault Management
  • Experience 3-4 years
  • Education Any
  • Work type On-site
  • Location United States
Check my match (free)
Google
AI Research Lab · 500+ Members · Mountain View, CA, United States

Google builds internet, software, cloud, and AI products used by consumers, developers, and organizations. Its portfolio includes Search, YouTube, Android, Chrome, Maps, Gmail, Workspace, Google Cloud, advertising platforms, devices, and Gemini AI products. The company develops large-scale computing infrastructure and research that power information retrieval, communication, productivity, media, navigation, and machine learning. Google is the largest operating business within Alphabet and earns a substantial share of its revenue from digital advertising.

All jobs at Google

130 more Software Engineer roles in Sunnyvale

Job Overview

Approx. salary range

207K – 314K

Our estimate — this employer did not publish a salary

Our estimate, not the employer’s. Worked out from the middle half of 84 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.

Eligibility
United States Right to work in the United States required.
Workplace
On-site
Job Posted:
2 weeks ago
Job Type
Full Time
Education
Any
Experience
3-4 years

Share This Job: