Senior Site Reliability Engineer
New
We sent a six-digit code to . Enter it below โ or use the link in the same email.
Or click the link in the same email โ either works.
Enter the email address on your account and we'll send you a link to set a new password.
Remembered it? Sign In
Looking for a site reliability engineer job in San Francisco? Choose from 34 open roles at 21 employers. Explore real salary data and skills breakdown across remote, hybrid and on-site roles.
Free. No credit card needed.
The roles on this page divide broadly into two patterns. Many are platform-oriented, asking engineers to own cloud infrastructure, manage Kubernetes clusters, and build automation with tools like Terraform and CI/CD pipelines. A notable subset is tied directly to AI or GPU workloads: for instance, the Site Reliability Engineer, Post Training at Thinking Machines Lab focuses on keeping large-scale distributed training systems reliable during live model runs, while the Senior SRE โ Fleet at Lambda centres on monitoring and automation for AI GPU clusters including InfiniBand and NCCL troubleshooting.
Others operate closer to product systems: the Senior SRE, Ads at Reddit covers ad-serving, auction, and billing reliability, and the Senior Infrastructure Engineer, SRE at Rocket Money involves establishing SLOs, error budgets, and disaster recovery for a high-transaction-volume platform. The Site Reliability Engineer, Production at Thinking Machines Lab combines CI/CD, observability, incident response, and multi-tenant isolation for a fine-tuning API. Titles vary across the sample โ "production engineering", "DevOps", and "infrastructure engineer" appear alongside the SRE label.
Kubernetes and Terraform are mentioned in the large majority of these postings, followed closely by Python, CI/CD pipelines, and AWS. GCP and GPU-related skills appear in roughly a third of postings, while Azure is less common. Docker appears in only a small minority despite being a foundational tool in many SRE stacks, which likely reflects description style rather than actual absence. These are mention counts across all postings on this page, not stated requirements.
The employers' own advertised ranges, for individual roles at different levels โ not an average and not a market rate.
Most postings on this page specify fully onsite work, making that the dominant arrangement in this sample. A smaller portion are listed as hybrid, and only a few are fully remote. No posting has an unknown work mode. Candidates who need flexibility should check individual listings carefully, as the ratio skews heavily toward in-person attendance โ the Lambda Fleet role, for example, states San Francisco or an alternative city as the base location. Hybrid terms, where offered, are not defined in the fact pack.
Jobs checked 49 minutes ago. ยท Guide reviewed 14 September 2026.
Search by role, company, or anything a posting mentions.