Together AI
Full TimeYou'll operate and optimize multi-petabyte storage systems for AI training workloads, managing parallel filesystems like Ceph and Weka, designing Kubernetes-native storage operators, and tuning end-to-end data paths for 10-50 GB/s per node. Requires 8+ years in distributed storage engineering at scale, production experience with GPU/HPC clusters, and strong Go and Python skills. Deep expertise in parallel filesystems, Kubernetes storage, and storage optimization for GPU workloads is essential. Remote-eligible role. $250,000–$300,000 base plus equity.
Written from this posting by Neural Jobs AI. The full description is below.
In this role, you will operate, scale, and optimize multi-petabyte storage systems purpose-built for the world’s largest AI training and inference workloads. You’ll manage and scale high-performance parallel filesystems and object stores, evaluate and integrate cutting-edge technologies such as Vast, Weka, Ceph, and Lustre, and solve the complex engineering challenges of operating at extreme throughput, low-latency data paths, and massive cluster-scale storage operations.
You will also build Kubernetes-native storage operators and self-service platforms that provide automated provisioning, strict multi-tenancy, performance isolation, and quota enforcement at cluster scale. Day-to-day, you’ll optimize end-to-end data paths for 10-50 GB/s per node, design multi-tier caching architectures, implement intelligent prefetching and model-weight distribution, and tune parallel filesystems for AI workloads.
Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers in our journey in building the next generation AI infrastructure.
We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $250,000 - $300,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.
Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
Please see our privacy policy at https://www.together.ai/privacy
Sign in and we will show how your role, experience, salary and location line up against what this employer asked for.
Check my match
Together AI is a cloud platform for training, fine-tuning and running open-source models, with a research arm contributing to open AI development.
Founded in 2022, Together AI provides high-performance inference for open models.
United States
Search by role, company, or anything a posting mentions.