logo
AI Resume Tailoring Sign in to use this AI Cover Letter Sign in to use this

Job Description

  • Develop scalable and sustainable system architecture and designs for products, services and enhancements.
  • Defend performance of critical services in alignment with customer expectations and SLOs.
  • Own and define strategy and set direction and establish roadmaps for Vertex AI services to increase reliability, efficiency and ultimately feature velocity.
  • Resolve outages or service disruptions and help design solutions to ensure systems are protected from similar classes of problems in the future.
  • Collaborate with development counterparts to incorporate and deliver enhancements to systems resulting in improved reliability, scalability and or performance.

Minimum qualifications:

  • Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.
  • 8 years of experience with software development in one or more programming languages.
  • 2 years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering roles managing large - scale infrastructure.
  • 3 years of experience designing, analyzing, and troubleshooting distributed systems.
  • 3 years of experience in machine learning infrastructure.

Preferred qualifications:

  • Master's degree in Computer Science or Engineering.
  • Experience enhancing and supporting large production systems on compute infrastructure.
  • Experience with networking, capacity and performance.
  • Experience in large-scale systems, architecture design and complex system integrations or migrations.
  • Experience supporting a Tier 1 rotation.
  • Expertise in SRE production principles and best practices.
  • Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.
  • 8 years of experience with software development in one or more programming languages.
  • 2 years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering roles managing large - scale infrastructure.
  • 3 years of experience designing, analyzing, and troubleshooting distributed systems.
  • 3 years of experience in machine learning infrastructure.
Do you match this job?

Here is what this employer asked for. Sign in and we will fill in your half.

  • Role Staff Software Engineer, Site Reliability Engineering, Vertex 1P GenAI SRE
  • Experience 8-9 years
  • Education Bachelor Degree, or equivalent experience
  • Work type On-site
  • Location United States
Check my match (free)
Google
AI Research Lab · 500+ Members · Mountain View, CA, United States

Google builds internet, software, cloud, and AI products used by consumers, developers, and organizations. Its portfolio includes Search, YouTube, Android, Chrome, Maps, Gmail, Workspace, Google Cloud, advertising platforms, devices, and Gemini AI products. The company develops large-scale computing infrastructure and research that power information retrieval, communication, productivity, media, navigation, and machine learning. Google is the largest operating business within Alphabet and earns a substantial share of its revenue from digital advertising.

All jobs at Google
Job Overview

Approx. salary range

214K – 311K

Our estimate — this employer did not publish a salary

Our estimate, not the employer’s. Worked out from the middle half of 59 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.

Eligibility
United States Right to work in the United States required.
Workplace
On-site
Job Posted:
2 days ago
Job Type
Full Time
Education
Bachelor Degree, or equivalent experience
Experience
8-9 years

Share This Job: