logo
AI Resume Tailoring Sign in to use this AI AI Cover Letter Sign in to use this AI

Job Description

Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.

Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.

About the role

You'll build Ollama’s cloud, a scalable inference platform that lets developers run large, capable open models in their workflow. You'll work on high-throughput, low-latency distributed systems — inference serving, GPU fleet management, routing, metering, and the platform that Pro, Max, Team, and Enterprise customers rely on to process trillions of tokens.

What you'll do

  • Build and scale the inference platform that serves every request from ollama.com.

  • Design the routing and capacity layer that places workloads across GPUs and regions for cost, latency, and availability.

  • Own multi-tenant infrastructure: isolation, quotas, usage metering, billing, and Pro/Max/team/enterprise tiering.

  • Build the reliability, observability, and cost controls for our team and customers

You may be a fit if

  • You have deep experience with high-throughput, low-latency distributed systems — inference serving, traffic routing, real-time data pipelines, or large-scale APIs.

  • You're comfortable with cost/performance tradeoffs at scale and have owned a production service end-to-end.

  • You've worked with Kubernetes, GPU scheduling, or inference infrastructure.

  • You think in terms of reliability, SLOs, and honest capacity planning.

  • Bonus: experience building an inference platform, GPU fleet management, or billing/metering for an AI service.

Do you match this job?

Sign in and we will show how your role, experience, salary and location line up against what this employer asked for.

Check my match
Ollama
Developer Tools & Platforms · 500+ Members · Palo Alto, CA, United States

Ollama develops software for running and managing open models. Its runtime lets developers download, execute, and interact with large language models locally through a command-line interface, local API, model library, and integrations with coding and AI tools. The product emphasizes simple setup, model choice, privacy, and control over where inference happens. Ollama has also expanded into hosted models and services, giving developers a consistent interface for using open-weight models on personal machines, servers, and supported cloud infrastructure.

All jobs at Ollama

Approx. salary range

205K – 338K

Our estimate — this employer did not publish a salary

Our estimate, not the employer’s. Worked out from the middle half of 25 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.

Eligibility

United States

Right to work in the United States required.

Workplace

On-site

Job Overview
Job Posted:
1 month ago
Job Type
Full Time
Education
Any
Experience
3-4 years

Share This Job: