logo

AI Job Summary

Lead AI-based solutions for data-center networking technologies, including agentic AI for management, predictive resilience, and optimization. This role requires a Ph.D. in electrical engineering, machine learning, computer science, or related field, plus 10+ years of experience in high-performance network architecture, data-center networking, or distributed systems. You'll need mastery of high-speed interconnect protocols like InfiniBand and advanced Ethernet, deep expertise in AI/ML and networking hardware, and proven success driving modern AI/ML projects. Location and compensation not specified.

Written from this posting by Neural Jobs AI. The full description is below.

AI Resume Tailoring Sign in to use this AI AI Cover Letter Sign in to use this AI

Job Description

We are seeking a highly skilled Principal Networking AI System Architect to join as a key contributor to the Applied Networking AI group. In this role you will scope and lead AI based solutions for networking technologies and drive their integration across teams.You’ll lead a portfolio and roadmap of projects that encompass agentic-AI for data-center management, predictive-resiliency, optimization and more. By collaborating closely with subject-matter-experts (SMEs), applied-researchers, product-managers, architects, data-engineers and other stakeholders you will push the envelope forward in using cutting-edge technologies and data-driven insights to improve NVIDIA's products.

What you'll be doing:

  • Build a shared roadmap and vision for AI based data-center management solutions spanning LLM intelligence for troubleshooting, predictive-resiliency and AIOPS, black-box optimization and performance tuning.
  • Work closely with engineering and reliability teams to scope and define workflows utilizing and benefitting from AI/ML.
  • Drive the integration of AI capabilities into system architecture and engineering workflows.
  • Identify system-level opportunities for failure management, automated troubleshooting, performance improvement, and resource optimization.
  • Translate system behavior, dependencies, data, and operational constraints into formulated research problems.

What we need to see:

  • Ph.D in electrical engineering, machine-learning, computer-science or another relevant field.
  • 10+ years of deep technical experience in high-performance network architecture, data center networking, or distributed systems design.
  • Mastery of high-speed interconnect protocols including InfiniBand and/or advanced Ethernet architectures.
  • Thorough experience driving high-impact projects centered on modern AI/ML such as LLMs/agents, deep-learning, black-box optimization or another relevant field.
  • Deep knowledge of AI/ML and networking-hardware/system-architecture.
  • Excellent ability to convey and communicate data-based insights to stakeholders and management.
  • Experience demonstrating an excellent track of collaboration with hands-on teams.

Ways to stand out from the crowd:

  • Demonstrated track record of architecting and deploying multi-thousand-node GPU clusters for hyperscale cloud environments.
  • Deep knowledge of NVIDIA networking technologies, including BlueField DPUs, Quantum InfiniBand switches, and Spectrum Ethernet platforms.
  • Expertise in in-network computing, telemetry, adaptive routing, and telemetry-driven network optimization.
  • High energy and a positive, proactive and curious approach.

Do you match this job?

Sign in and we will show how your role, experience, salary and location line up against what this employer asked for.

Check my match
NVIDIA
Hardware & Semiconductors · 200-500 Members · Santa Clara, CA, United States

NVIDIA builds the GPUs and the CUDA software stack that most modern AI is trained and served on, along with its own research in graphics, robotics and foundation models.

All jobs at NVIDIA

Location

Israel

Job Overview
Job Posted:
19 hours ago
Workplace
On-site
Job Type
Full Time
Education
Any
Experience
8+ years

Share This Job: