logo

AI Job Summary

Optimize AI inference performance across application, model, and infrastructure layers to increase throughput-per-GPU and reduce latency. Requires a bachelor's degree in a technical field or equivalent experience, eight years of software development, and proficiency in Python and C++. Must understand AI model execution constraints, throughput-latency tradeoffs, and modern serving architectures. Experience with LLM inference serving, ML profilers, and GPU performance concepts preferred.

Written from this posting by Neural Jobs AI. The full description is below.

AI Resume Tailoring Sign in to use this AI Cover Letter Sign in to use this

Job Description

  • Analyze and optimize AI inference workloads across the application, model, and distributed fleet infrastructure layers to methodically increase throughput-per-GPU and reduce latency.
  • Design and implement inference optimization techniques.
  • Investigate and resolve complex model inference performance bottlenecks across the stack.
  • Model the latency-to-cost impacts of system variables (such as batch-sizing and utilization goals) and translate these insights into actionable signals that drive production systems.
  • Develop investigative tools and metrics (e.g., compute/FLOPs funnels) that track where compute is spent across the fleet.

Minimum qualifications:

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related technical field, or equivalent practical experience.
  • 8 years of experience in software development.
  • Experience in Python and C++, including navigating, debugging, and modifying serving codebases.
  • Experience with AI model execution constraints, throughput-latency tradeoffs, memory bandwidth limitations, and modern serving architectures.

Preferred qualifications:

  • Experience with real world LLM inference serving environments or direct contributions to modern open-source inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang, Dynamo).
  • Experience profiling workloads using standard ML profilers (e.g., PyTorch profiler) and internal trace analysis tools.
  • Experience with observability and reliability for large distributed systems.
  • Familiarity with GPU/TPU/accelerator performance concepts (e.g. memory bandwidth, quantization, collective communication, kernel), and can reason their implications to the overall inference serving performance.
  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related technical field, or equivalent practical experience.
  • 8 years of experience in software development.
  • Experience in Python and C++, including navigating, debugging, and modifying serving codebases.
  • Experience with AI model execution constraints, throughput-latency tradeoffs, memory bandwidth limitations, and modern serving architectures.
Do you match this job?

Here is what this employer asked for. Sign in and we will fill in your half.

  • Role Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind
  • Experience 8-9 years
  • Education Bachelor Degree, or equivalent experience
  • Work type On-site
  • Location United States
Check my match (free)
Google
AI Research Lab · 500+ Members · Mountain View, CA, United States

Google builds internet, software, cloud, and AI products used by consumers, developers, and organizations. Its portfolio includes Search, YouTube, Android, Chrome, Maps, Gmail, Workspace, Google Cloud, advertising platforms, devices, and Gemini AI products. The company develops large-scale computing infrastructure and research that power information retrieval, communication, productivity, media, navigation, and machine learning. Google is the largest operating business within Alphabet and earns a substantial share of its revenue from digital advertising.

All jobs at Google

145 more Software Engineer roles in Mountain View

Job Overview

Approx. salary range

209K – 286K

Our estimate — this employer did not publish a salary

Our estimate, not the employer’s. Worked out from the middle half of 44 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.

Eligibility
United States Right to work in the United States required.
Workplace
On-site
Job Posted:
3 days ago
Job Type
Full Time
Education
Bachelor Degree, or equivalent experience
Experience
8-9 years

Share This Job: