Senior Network Site Reliability Engineer
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.
SpaceXAI is looking for a highly skilled and versatile OT (Operational Technology) Systems Engineer to implement and support the backend infrastructure for next-generation controls and industrial software platforms critical to our hyperscale AI supercomputer campuses and co-located power generation. This role sits with the Supercomputer Physical Infrastructure team and supports dedicated controls, facilities, and power organizations while leveraging adjacent IT expertise, tooling, and technologies. You will balance sustainment of live cooling, power, and facility control systems with modernization and continuous improvement. The ideal candidate thrives in high-stakes, 24/7 environments, brings a strong sense of urgency balanced with operational excellence, and combines deep OT expertise with IT technologies such as virtualization, VDI, GitOps, network-level redundancy, and edge compute to create streamlined and secure ICS environments for ultra-dense AI compute.
SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.
Sign in and we will show how your role, experience, salary and location line up against what this employer asked for.
Check my match
xAI develops Grok, a family of large language models, and builds the large-scale training and inference infrastructure behind them, including the Colossus supercomputing cluster.
Founded in 2023, the company focuses on frontier model research and deployment across the X platform and its own products.
246K – 344K
Our estimate — this employer did not publish a salaryOur estimate, not the employer’s. Worked out from the middle half of 23 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.
United States
Right to work in the United States required.On-site
Search by role, company, or anything a posting mentions.