Site Reliability Engineer Jobs in San Francisco, USA

Looking for a site reliability engineer job in San Francisco? Choose from 34 open roles at 21 employers. Explore real salary data and skills breakdown across remote, hybrid and on-site roles.

34 Open roles
21 Companies hiring
258K Average salary

What do Site reliability engineers in San Francisco work on?

The roles on this page divide broadly into two patterns. Many are platform-oriented, asking engineers to own cloud infrastructure, manage Kubernetes clusters, and build automation with tools like Terraform and CI/CD pipelines. A notable subset is tied directly to AI or GPU workloads: for instance, the Site Reliability Engineer, Post Training at Thinking Machines Lab focuses on keeping large-scale distributed training systems reliable during live model runs, while the Senior SRE โ€“ Fleet at Lambda centres on monitoring and automation for AI GPU clusters including InfiniBand and NCCL troubleshooting.

Others operate closer to product systems: the Senior SRE, Ads at Reddit covers ad-serving, auction, and billing reliability, and the Senior Infrastructure Engineer, SRE at Rocket Money involves establishing SLOs, error budgets, and disaster recovery for a high-transaction-volume platform. The Site Reliability Engineer, Production at Thinking Machines Lab combines CI/CD, observability, incident response, and multi-tenant isolation for a fine-tuning API. Titles vary across the sample โ€” "production engineering", "DevOps", and "infrastructure engineer" appear alongside the SRE label.

Skills and experience employers ask for

Kubernetes and Terraform are mentioned in the large majority of these postings, followed closely by Python, CI/CD pipelines, and AWS. GCP and GPU-related skills appear in roughly a third of postings, while Azure is less common. Docker appears in only a small minority despite being a foundational tool in many SRE stacks, which likely reflects description style rather than actual absence. These are mention counts across all postings on this page, not stated requirements.

Kubernetes 74%
Terraform 68%
AWS 62%
CI/CD 62%
Python 59%
GCP 38%
Agents 26%
Azure 26%
GPUs 24%
APIs 21%
LLMs 18%
Docker 15%
TypeScript 15%
Evaluation 12%

Site Reliability Engineer roles that state a salary

The employers' own advertised ranges, for individual roles at different levels โ€” not an average and not a market rate.

San Francisco office, hybrid or remote?

Most postings on this page specify fully onsite work, making that the dominant arrangement in this sample. A smaller portion are listed as hybrid, and only a few are fully remote. No posting has an unknown work mode. Candidates who need flexibility should check individual listings carefully, as the ratio skews heavily toward in-person attendance โ€” the Lambda Fleet role, for example, states San Francisco or an alternative city as the base location. Hybrid terms, where offered, are not defined in the fact pack.

Make your application specific to the work

  1. Read the full job description SRE postings here vary significantly in specialisation โ€” GPU infrastructure, ad systems, AI training pipelines โ€” so check the specific scope before applying to avoid mismatched expectations on either side.
  2. Map your experience to the role level Seniority ranges from unspecified through senior, staff, lead, and manager; each tier typically carries different ownership expectations around incident response, architecture, and cross-team influence.
  3. Prepare your infrastructure examples Most descriptions ask for demonstrated experience with distributed systems, cloud platforms, and automation tooling, so have concrete examples ready covering observability, incident handling, and reliability improvements you have owned.
  4. Note visa and degree requirements per posting Sponsorship eligibility and degree requirements are stated at the individual posting level and are not uniform across this sample โ€” the Thinking Machines Lab production role explicitly states a bachelor's degree or equivalent, but other postings vary.

Questions about Site Reliability Engineer jobs in San Francisco

Pay is advertised in a minority of these postings. The ranges vary widely by employer and seniority: individual advertised bands run from a lower bound around 165,000 up to highs above 390,000 in some postings. These figures are from specific advertised ranges on this page only and are not representative of all employers listed here, most of whom have not published a pay band.

A small number of postings on this page are listed as remote. The majority are onsite, with a portion hybrid. Where remote or hybrid working is important to you, you should confirm the specific terms with each employer, as the fact pack does not detail hybrid schedules.

Degree requirements differ by posting. The Thinking Machines Lab production engineering role states a bachelor's degree in computer science or equivalent experience. Other postings do not specify, so this should be checked individually rather than assumed.

Kubernetes and Terraform are mentioned in the large majority of postings on this page. Python, CI/CD, and AWS follow closely. GCP and GPU expertise appear in around a third of postings. These are skill mentions in job descriptions, not uniform requirements across all roles.

It depends on the employer. Several postings โ€” particularly from Lambda and Thinking Machines Lab โ€” specifically involve GPU clusters, AI training pipelines, or LLM-related infrastructure. Others, such as the Reddit Ads and Rocket Money roles, focus on more traditional product and platform reliability without an AI infrastructure focus.

Jobs checked 49 minutes ago. ยท Guide reviewed 14 September 2026.