Design, build, and operate production datacenter and core networks at scale, handling routing and switching configuration, platform qualification, incident troubleshooting, and automation. Requires several years of hands-on experience designing or operating production networks with solid BGP and interior routing protocol knowledge, TCP/IP, VLANs, and Ethernet expertise. Python or Ansible automation and datacenter vendor experience preferred. Onsite in Palo Alto with travel to build sites; $150,000–$250,000 base salary.
Written from this posting by Neural Jobs AI. The full description is below.
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.
ABOUT THE ROLE:
SpacexAI is building and operating large-scale networks that underpin training and inference infrastructure, including high-performance / supercompute fabrics that connect GPU clusters, plus the core, edge, and datacenter networks that keep that infrastructure reachable and reliable.
We need a Network Engineer who is strong on fundamentals and comfortable owning production network design, deployment, and operations end to end. This is a hands-on engineering seat — not a NOC technician role and not a network-software (telemetry/ZTP platform) SWE role. You will design and build networks, qualify platforms, ship changes safely, and keep availability and performance high as we scale.
Travel to Memphis (and other build sites) may be required for capacity build-outs. You will participate in a team on-call rotation.
RESPONSIBILITIES:
Design, deploy, and operate production datacenter and campus/core networks at scale
Own routing and switching configuration standards (BGP and at least one IGP such as OSPF or IS-IS), including change design, peer reviews, and execution
Qualify new network platforms, optics, and topologies; contribute to architecture and capacity planning
Build and improve monitoring, alerting, and operational documentation so issues are caught and fixed quickly
Troubleshoot Layer 2/Layer 3 incidents end to end — from link flaps and optics through routing and traffic engineering — and drive root cause and lasting fixes
Automate repetitive network tasks with Python, Ansible, or similar tooling where it reduces toil
Partner with compute, facilities, and software teams during cluster build-outs and maintenance windows
Support high-performance / supercompute network environments (Ethernet AI/HPC fabrics, RoCE/RDMA-capable designs) as part of the broader network estate — deep specialist RoCE/NCCL ownership is a plus, not the bar for this seat
BASIC QUALIFICATIONS:
Several years designing and/or operating production networks in a datacenter, ISP, cloud, or large enterprise environment
Solid hands-on experience with BGP and at least one interior routing protocol
Working knowledge of TCP/IP, VLANs, EVPN/VXLAN or equivalent datacenter overlays, and optics / high-speed Ethernet
Experience troubleshooting live production network incidents and participating in on-call
Strong written and verbal communication; clear change docs and incident notes
PREFERRED SKILLS AND EXPERIENCE:
Experience with modern datacenter vendors (e.g. Arista, Cisco, Juniper, Nvidia/Mellanox)
Familiarity with high-performance or supercompute networking (RoCEv2, congestion control, GPU cluster fabrics) — useful context for our environment, not a hard filter
Network automation (Python, Ansible, Terraform, or similar) used in production
Experience with EVPN, leaf-spine, and large-scale Ethernet fabrics
Prior work supporting rapid datacenter or cluster capacity build-outs
ADDITIONAL REQUIREMENTS:
Willing to work onsite in Palo Alto
COMPENSATION AND BENEFITS:
$150,000 - $250,000 USD
Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.
SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.
Do you match this job?
Here is what this employer asked for. Sign in and we will fill in your half.
Foundation Models · 500+ Members · Palo Alto, CA, United States
xAI develops Grok, a family of large language models, and builds the large-scale training and inference infrastructure behind them, including the Colossus supercomputing cluster.
Founded in 2023, the company focuses on frontier model research and deployment across the X platform and its own products.
Our estimate — this employer did not publish a salary
Our estimate, not the employer’s. Worked out from the middle half of 35 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.
Eligibility
United States
Right to work in the United States required.