logo
AI Resume Tailoring Sign in to use this AI AI Cover Letter Sign in to use this AI

Job Description

  • Collaborate closely with Research teams to understand next generation modeling approaches, ensuring they are designed and implemented with production considerations in mind.
  • Work with infrastructure teams to deliver serving infrastructure that is designed for maximum efficiency and performance, addressing bottlenecks in speed, scale, and quality.
  • Identify opportunities to automate tasks, eliminate redundancies, build performant tests, and improve the overall velocity of model releases.
  • Gain a deep understanding of serving frameworks, pre-processing pipelines, caching mechanisms, and other relevant technologies.
  • Leverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers (XLA), custom kernels (Pallas), and serving infrastructure on hardware accelerators (TPUs/GPUs).

Minimum qualifications:

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 2 years of experience in deploying and maintaining machine learning models in a live production environment.
  • Experience in profiling, configuring, or executing ML workloads directly on hardware accelerators (e.g., GPU or TPU).
  • Experience designing, building, or optimizing model serving infrastructure or inference backends.

Preferred qualifications:

  • Experience with developing serving infrastructure.
  • Experience programming hardware accelerators (GPUs, TPUs) via ML frameworks (e.g., JAX, PyTorch) or low-level programming models (e.g., Pallas, CUDA, OpenCL).
  • Experience profiling software to identify performance bottlenecks.
  • Experience with distributed ML systems optimization and parallelism (e.g., data, model, or pipeline parallelism).
  • Familiarity with writing performance-optimized kernels.
  • Understanding of LLM architecture and inference performance dynamics (e.g., Transformer models, memory bandwidth and compute bounds, KV cache scaling).
  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 2 years of experience in deploying and maintaining machine learning models in a live production environment.
  • Experience in profiling, configuring, or executing ML workloads directly on hardware accelerators (e.g., GPU or TPU).
  • Experience designing, building, or optimizing model serving infrastructure or inference backends.
Do you match this job?

Sign in and we will show how your role, experience, salary and location line up against what this employer asked for.

Check my match
Google
AI Research Lab · 500+ Members · Mountain View, CA, United States

Google builds internet, software, cloud, and AI products used by consumers, developers, and organizations. Its portfolio includes Search, YouTube, Android, Chrome, Maps, Gmail, Workspace, Google Cloud, advertising platforms, devices, and Gemini AI products. The company develops large-scale computing infrastructure and research that power information retrieval, communication, productivity, media, navigation, and machine learning. Google is the largest operating business within Alphabet and earns a substantial share of its revenue from digital advertising.

All jobs at Google

Location

United Kingdom

Job Overview
Job Posted:
3 weeks ago
Workplace
On-site
Job Type
Full Time
Education
Any
Experience
3+ years

Share This Job: