DevJobs

Principal Networking AI Systems Architect

Overview
Skills
  • Deep learning Deep learning
  • ML ML
  • AI
  • Black-box optimization
  • Data Center Networking
  • Distributed Systems
  • Ethernet
  • InfiniBand
  • LLM
  • Telemetry-driven network optimization
  • Adaptive routing
  • Telemetry
  • Spectrum Ethernet platforms
  • Quantum InfiniBand switches
  • NVIDIA networking technologies
  • In-network computing
  • BlueField DPUs
We are seeking a highly skilled Principal Networking AI System Architect to join as a key contributor to the Applied Networking AI group. In this role you will scope and lead AI based solutions for networking technologies and drive their integration across teams.You’ll lead a portfolio and roadmap of projects that encompass agentic-AI for data-center management, predictive-resiliency, optimization and more. By collaborating closely with subject-matter-experts (SMEs), applied-researchers, product-managers, architects, data-engineers and other stakeholders you will push the envelope forward in using cutting-edge technologies and data-driven insights to improve NVIDIA's products.

What You'll Be Doing

  • Build a shared roadmap and vision for AI based data-center management solutions spanning LLM intelligence for troubleshooting, predictive-resiliency and AIOPS, black-box optimization and performance tuning.
  • Work closely with engineering and reliability teams to scope and define workflows utilizing and benefitting from AI/ML.
  • Drive the integration of AI capabilities into system architecture and engineering workflows.
  • Identify system-level opportunities for failure management, automated troubleshooting, performance improvement, and resource optimization.
  • Translate system behavior, dependencies, data, and operational constraints into formulated research problems.

What We Need To See

  • Ph.D in electrical engineering, machine-learning, computer-science or another relevant field.
  • 10+ years of deep technical experience in high-performance network architecture, data center networking, or distributed systems design.
  • Mastery of high-speed interconnect protocols including InfiniBand and/or advanced Ethernet architectures.
  • Thorough experience driving high-impact projects centered on modern AI/ML such as LLMs/agents, deep-learning, black-box optimization or another relevant field.
  • Deep knowledge of AI/ML and networking-hardware/system-architecture.
  • Excellent ability to convey and communicate data-based insights to stakeholders and management.
  • Experience demonstrating an excellent track of collaboration with hands-on teams.

Ways To Stand Out From The Crowd

  • Demonstrated track record of architecting and deploying multi-thousand-node GPU clusters for hyperscale cloud environments.
  • Deep knowledge of NVIDIA networking technologies, including BlueField DPUs, Quantum InfiniBand switches, and Spectrum Ethernet platforms.
  • Expertise in in-network computing, telemetry, adaptive routing, and telemetry-driven network optimization.
  • High energy and a positive, proactive and curious approach.

, , JR2024255

Nvidia