Machine Learning Engineer, Platform
New
We sent a six-digit code to . Enter it below — or use the link in the same email.
Or click the link in the same email — either works.
Enter the email address on your account and we'll send you a link to set a new password.
Remembered it? Sign In
Graphcore
NewBuild and lead a new SRE organization from inception through production launch for a rapidly scaling AI supercomputing platform. You'll establish the operating model, hire and mentor the team, embed reliability practices into platform development, and lead 24x7x365 production operations. This role requires significant SRE or infrastructure reliability leadership experience, strong incident management skills, hands-on Linux and distributed systems knowledge, and the ability to work at both strategic and technical levels. The position is based in the UK with flexible working options.
Written from this posting by Neural Jobs AI. The full description is below.
About Graphcore
How often do you get the chance to build a technology that transforms the future of humanity?
Graphcore products have set the standard in made-for-AI compute hardware and software, gaining global attention and industry acclaim. Now we are developing the next generation of artificial intelligence compute with systems that will allow AI researchers to develop more advanced models, help scientists unlock exciting new discoveries, and power companies around the world as they put AI at the heart of their business.
Graphcore recently joined SoftBank Group – bringing large and ongoing investment from one of the world’s leading backers of innovative AI companies.
Job Summary
We are seeking an experienced Site Reliability Engineering leader to build and lead a new SRE organization responsible for the production operation of a rapidly scaling AI supercomputing platform. The environment combines highly customized compute, high-performance networking, storage and supporting infrastructure, and will grow through multiple phases of deployment.
This is a rare opportunity to establish the reliability function for a new platform from the ground up. The platform and its operational model are being developed in parallel and will ultimately support a 24x7x365 production service with stringent availability requirements.
You will take the SRE organization from initial formation through production launch, stabilization and scale. This includes hiring and developing the team, defining the operating model, establishing production readiness and incident-management practices, and ensuring reliability and operability are engineered into the platform from the outset.
SRE is responsible for the operational capability required to run the platform reliably in production, while partnering with engineering teams that remain accountable for the reliability and operability of the systems they build.
This is not a purely managerial position. During the development and early production phases, the SRE Manager will be expected to work directly with engineering teams, develop a deep understanding of the platform, and participate in troubleshooting and incident response.
Over time, success will increasingly mean building the people, processes, automation, tooling, and operational discipline that allow the organization to operate effectively without depending on you for day-to-day escalation.
Responsibilities and Duties
Required Skills and Experience
Desired but Not Required
Candidates are not expected to have experience in all the areas below. Experience in several would be particularly valuable:
What Success Looks Like
In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.
Here is what this employer asked for. Sign in and we will fill in your half.
Graphcore designs the Intelligence Processing Unit, a processor architected specifically for machine-intelligence workloads, together with the Poplar software stack that compiles models onto it.
Founded in Bristol in 2016 and acquired by SoftBank in 2024, the company builds silicon and systems for AI training and inference.
197K – 299K
Our estimate — this employer did not publish a salaryOur estimate, not the employer’s. Worked out from the middle half of 62 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.
Already have an account? Sign in
Continue without an account and apply on the Graphcore website
Search by role, company, or anything a posting mentions.