logo

AI Job Summary

You'll optimize ML systems for throughput and latency at scale, working with PyTorch, inference engines like vLLM or TensorRT, and Nvidia GPUs. This role requires 5+ years building high-performance code and direct experience improving GPU utilization—debugging occupancy issues, optimizing algorithms, or reducing host overhead. You'll contribute to Modal's container runtime and open-source projects. Knowledge of Linux kernel and container internals is a bonus.

Written from this posting by Neural Jobs AI. The full description is below.

AI Resume Tailoring Sign in to use this AI AI Cover Letter Sign in to use this AI

Job Description

About Us:

AI needs a new infrastructure layer. We're building it at Modal.

Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now.

Our customers include category-defining companies like Lovable, Ramp, Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale.

We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September.

Our team includes creators of popular open-source projects (e.g.,Seaborn,Luigi), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience.

The Role:

We are looking for strong engineers with experience in making ML systems performant at scale. If you are interested in contributing to open-source projects and Modal’s container runtime to push language and diffusion models towards higher throughput and lower latency, we’d love to hear from you!

Requirements:

  • 5+ years of experience writing high-quality, high-performance code.

  • Experience working with torch, high-level ML frameworks, and inference engines (vLLM or TensorRT).

  • Familiarity with Nvidia GPU architecture and CUDA.

  • Experience with ML performance engineering (tell us a story about boosting GPU performance — debugging SM occupancy issues, rewriting an algorithm to be compute-bound, eliminating host overhead, etc).

  • Nice-to-have: familiarity with low-level operating system foundations (Linux kernel, file systems, containers, etc).

Do you match this job?

Sign in and we will show how your role, experience, salary and location line up against what this employer asked for.

Check my match
Modal Labs
AI Infrastructure & Compute · 50-100 Members · New York, NY, United States

Modal is a serverless compute platform for AI and data workloads, letting developers run models, batch jobs and GPU tasks from Python without managing infrastructure.

Founded in 2021 in New York, Modal is used for fine-tuning, inference and large-scale data processing.

All jobs at Modal Labs

Salary

200K - 350K Yearly

Location

United States

Job Overview
Job Posted:
4 months ago
Workplace
On-site
Job Type
Full Time
Education
Any
Experience
8+ years

Share This Job: