Together AI
Full TimeBuild provisioning state machines and self-service APIs that automatically manage GPU cluster infrastructure from bare metal to production. You'll design declarative workflows and control planes using tools like Temporal or Cadence, implement event-driven reconciliation systems similar to Kubernetes operators, and own end-to-end reliability in production. Required: strong software engineering in Go, Python, or Rust; experience with workflow orchestration and state reconciliation systems; and familiarity with event-driven architectures. All experience levels considered.
Written from this posting by Neural Jobs AI. The full description is below.
We're looking for a Software Engineer to build the systems that treat infrastructure as software. This role owns the software state machines that provision hardware, bring it into service, and manage its full lifecycle — turning racks of GPUs into running inference clusters without a human touching a runbook. The Research and Inference team is your customer: today they file tickets and wait; the target state is that they issue a single API call to stand up, scale, or tear down a cluster, and the system takes care of the rest. The platform is manifest-driven such that teams declare the desired state of a cluster or host — shape, topology, software stack — and the system is responsible for reconciling reality to that manifest, continuously, through every stage of its lifecycle. You will design the engines that manifest the schema, the engines that execute against it, and the workflows that carry a piece of hardware or a cluster from one state to the next—taking it from bare metal to a fully functioning AI cluster for training or inference.
You'll write production code which is typed, tested, versioned, and deployed through CI/CD that models infrastructure state and reconciles it, the same way a Kubernetes controller reconciles a cluster's desired state. Success looks like eliminating manual provisioning work, not documenting it better.
A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship.You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production.
Core requirements (all levels):
Nice to have:
Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers and engineers in our journey in building the next generation AI infrastructure.
Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
Please see our privacy policy at https://www.together.ai/privacy.
Sign in and we will show how your role, experience, salary and location line up against what this employer asked for.
Check my match
Together AI is a cloud platform for training, fine-tuning and running open-source models, with a research arm contributing to open AI development.
Founded in 2022, Together AI provides high-performance inference for open models.
Netherlands
Search by role, company, or anything a posting mentions.