logo

AI Job Summary

Design and scale inference and model-serving infrastructure from ground up through production deployment, optimizing for latency, throughput, and reliability at high concurrency. This role requires five or more years building ML inference systems in production, hands-on experience with tools like TensorFlow Serving, TorchServe, Triton, or KServe, and strong distributed systems knowledge with Docker and Kubernetes. On-site in San Mateo, California.

Written from this posting by Neural Jobs AI. The full description is below.

AI Resume Tailoring Sign in to use this AI Cover Letter Sign in to use this

Job Description

About the Role

This is a hands-on ML Infrastructure Engineer role at an early-stage enterprise AI company building a context and data governance layer that makes AI agents reliable in production. You will own the inference and model-serving infrastructure end to end, ensuring agents run fast and reliably at increasing concurrency. The work is squarely production-focused with real-world impact across regulated industries like insurance, banking, healthcare, and asset management.

What You'll Do

  • Design, build, and scale inference and model-serving infrastructure from the ground up through production deployment.

  • Optimize systems for latency, throughput, and reliability under high concurrency.

  • Collaborate closely with ML and infrastructure teams to ensure seamless integration and surface performance bottlenecks.

  • Drive solutions to infrastructure challenges across a fast-moving, cross-functional team.

What We're Looking For

  • 5 or more years building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments.

  • Hands-on experience designing and scaling inference-serving systems using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom solutions.

  • Strong distributed systems fundamentals, including containerization and orchestration with Docker and Kubernetes.

  • Proficiency with monitoring and observability tooling for production systems, such as Prometheus, Grafana, or distributed tracing frameworks.

  • Experience deploying and managing ML workloads on cloud platforms (AWS, GCP, or Azure).

  • Proficiency in at least one systems or backend language: Python, Go, Rust, C++, or Java.

  • Comfort collaborating across both ML and infrastructure disciplines in a fast-paced environment.

  • Nice to have: experience with knowledge graphs, semantic search, or graph databases; real-time or low-latency inference systems; agentic or multi-step AI pipelines; enterprise data integration or pipeline infrastructure.

Location

On-site in San Mateo, California, United States. Visa sponsorship is not available for this role.

Do you match this job?

Here is what this employer asked for. Sign in and we will fill in your half.

  • Role ML Infrastructure Engineer
  • Experience 3-4 years
  • Work type On-site
  • Location United States
Check my match (free)
Clera
AI & Machine Learning · 500+ Members · United States

Clera is an AI recruiting platform that introduces candidates directly to hiring managers at the companies they want to work for.

All jobs at Clera

22 more Machine Learning Engineer roles in San Mateo

Job Overview

Approx. salary range

201K – 310K

Our estimate — this employer did not publish a salary

Our estimate, not the employer’s. Worked out from the middle half of 98 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.

Eligibility
United States Right to work in the United States required.
Workplace
On-site
Job Posted:
2 days ago
Job Type
Full Time
Experience
3-4 years

Share This Job: