logo

AI Job Summary

You'll own reliability across a global GPU inference platform, building Kubernetes infrastructure and automation to handle hardware failures, traffic spikes, and multi-region operations. The role requires production experience with distributed systems, strong Linux and networking fundamentals, and hands-on Kubernetes expertise. You'll write software to automate operations, improve observability, respond to incidents, and work directly with infrastructure and platform engineers in a flat organization. The position is based on-site or remote; no salary was specified.

Written from this posting by Neural Jobs AI. The full description is below.

AI Resume Tailoring Sign in to use this AI Cover Letter Sign in to use this

Job Description

AI companies need inference that’s fast, reliable, and economical at scale. Parasail delivers it. We’re building an enterprise-grade inference cloud for open-weight models where customers pay for the tokens they use, and we handle everything required to serve them.

Behind one OpenAI-compatible API, we pool GPU capacity from providers around the world and continuously optimize where and how workloads run. That means turning a changing mix of hardware, networks, and infrastructure into a service customers can trust.

We’ve raised a $32 million Series A, and we’re scaling beyond trillions of tokens a day. You’ll have the ownership and reach to shape how we get there.

The Role

At Parasail, reliability is an engineering problem that spans the entire stack. A GPU fails. A provider goes down. Traffic spikes. Customers still expect their inference to work.

We’re hiring Site Reliability Engineers to build the systems that make that possible. You’ll own infrastructure across our global GPU fleet, write software that automates operations, and make the platform better at detecting, surviving, and recovering from failures.

You’ll work directly with infrastructure, platform, and inference engineers in a flat organization. We welcome SREs, software engineers, platform engineers, and systems engineers who want to build ambitious systems and take responsibility for how they perform in production.

What You’ll Do

  • Scale a global GPU fleet. Build and improve the Kubernetes infrastructure behind provisioning, networking, storage, and service deployment across providers and regions.

  • Make failure survivable. Design better isolation, failover, and recovery so hardware and infrastructure failures have less impact on customers.

  • Build software that runs infrastructure. Automate capacity expansion, deployments, and maintenance, eliminating manual work and making changes safer.

  • Make the system understandable. Develop observability and diagnostics that reveal bottlenecks, surface failures, and help engineers act quickly.

  • Own the production feedback loop. Respond to incidents, get to the root cause, and turn what you learn into stronger systems.

  • Push the platform forward. Work across the stack to improve performance, utilization, security, and reliability as inference demand grows.

What You Bring

  • Experience building and operating production infrastructure or distributed systems, with real ownership of reliability.

  • Strong Linux fundamentals and practical knowledge of networking, storage, and containers.

  • Hands-on experience running Kubernetes in production.

  • The ability to write maintainable software and automation to solve infrastructure problems.

  • A systematic approach to debugging problems that cross application, cluster, network, and hardware boundaries.

  • Good judgment about when to move quickly, when to simplify, and where reliability matters most.

  • The initiative to take a problem from investigation through implementation and work closely with teammates along the way.

Your strongest skill might be software development, distributed systems, or infrastructure operations. We’re building a team with complementary strengths; your previous job title matters less than what you can build and own.

Nice to Have

  • Experience with multi-region, multi-provider, or bare-metal infrastructure.

  • Familiarity with GPUs, model serving, or inference systems such as vLLM or SGLang.

  • Experience with infrastructure as code, CI/CD, observability, or automated recovery.

  • Experience building highly available services, multi-tenant platforms, or distributed data systems.

Why Join Parasail

The systems you build will determine how reliably and efficiently customers can run AI in production. You’ll work close to the hardware, deep in distributed systems, and alongside engineers optimizing the inference stack.

This is a small team tackling problems at substantial scale. You’ll own meaningful architecture decisions, ship improvements directly into production, and help build the foundation for the next stage of AI infrastructure.

Do you match this job?

Here is what this employer asked for. Sign in and we will fill in your half.

  • Role Senior Site Reliability Engineer
  • Experience 5-7 years
  • Work type On-site
  • Location United States
Check my match (free)
Parasail
AI Infrastructure & Compute · United States

Parasail is the inference cloud for AI-native startups. Run any open model with production reliability, flexible scaling, and per-token pricing.

All jobs at Parasail
Job Overview

Approx. salary range

188K – 292K

Our estimate — this employer did not publish a salary

Our estimate, not the employer’s. Worked out from the middle half of 90 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.

Eligibility
United States Right to work in the United States required.
Workplace
On-site
Job Posted:
2 days ago
Job Type
Full Time
Experience
5-7 years

Share This Job: