Search

Lead GPU Cluster Solution Architect

PublishedPublished: 6/14/2022
Technology

Job Description

ABOUT THE ROLE

\n

Axe Compute is seeking a Lead Architect, GPU Cluster Solutions to design GPU cluster configurations for prospective and signed engagements, translating client requirements and NVIDIA Reference Architecture into a buildable, supportable design, spanning compute, storage, networking, software and spares strategy for support.

\n


\n

ROLE AT A GLANCE

\n

    \n
  • Mandate: Own end-to-end technical design of GPU cluster deployments, from client requirements through NVIDIA Reference Architecture compliance and sparing strategy.
  • \n

  • Scope: Cluster design, network architecture (InfiniBand/RoCE/Ethernet), sparing and spares planning, connectivity design (internet/VPN/firewall, dedicated circuits), design adjustments for site and hardware constraints.
  • \n

  • Key Outcomes: Designs that meet client SLAs and NVIDIA Reference Architecture standards, sparing plans that protect uptime commitments, designs that account for real-world site and hardware lead-time constraints.
  • \n

\n


\n

WHAT YOU WILL OWN

\n

Cluster Design & Reference Architecture

\n

    \n
  • Design GPU cluster configurations (compute, storage, networking) against NVIDIA Reference Architecture for each signed engagement.
  • \n

  • Translate client technical requirements into a complete bill of design, including all necessary compute, storage, and networking components.
  • \n

\n

Network Architecture

\n

    \n
  • Design network topology and fabric selection, including InfiniBand, RoCE, and Ethernet options, appropriate to each client's workload and performance requirements.
  • \n

  • Incorporate internet, VPN, and firewall connectivity requirements into cluster designs.
  • \n

  • Design dedicated point-to-point network requirements where needed, including protected optical circuits and similar dedicated connectivity.
  • \n

\n

Sparing & Availability Strategy

\n

    \n
  • Formulate and own the hot/cold sparing plan for each deployment to meet contracted SLA commitments.
  • \n

  • Adjust sparing and design assumptions based on data center power/cooling parameters and hardware lead-time constraints.
  • \n

\n

Design Adaptation & Site Constraints

\n

    \n
  • Adjust cluster designs to fit site-specific power, cooling, and space constraints identified by the Data Center Procurement and Operations Director.
  • \n

  • Work with Supply Chain to align design decisions with realistic hardware delivery timing.
  • \n

\n

Cross-Functional Collaboration

\n

    \n
  • Partner with the VP, Deployments and Deployment Program Manager to ensure designs translate cleanly into buildable, trackable project plans.
  • \n

  • Support acceptance test design and criteria definition, ensuring test procedures validate the as-designed architecture.
  • \n

\n


\n

REQUIRED QUALIFICATIONS

\n

    \n
  • 7+ years in solutions architecture, network engineering, or systems engineering supporting GPU, HPC, or large-scale compute infrastructure.
  • \n

  • Deep working knowledge of NVIDIA Reference Architecture (HGX, NVL72) and GPU cluster design principles.
  • \n

  • Hands-on experience with InfiniBand, RoCE, and high-speed Ethernet fabric design.
  • \n

  • Experience designing sparing/spares strategies for mission-critical infrastructure.
  • \n

  • Experience incorporating firewall, VPN, and dedicated circuit (e.g., protected optical) requirements into network designs.
  • \n

  • Experience with high-speed shared storage solutions (e.g., Weka, Vast, DDN).
  • \n

\n


\n

PREFERRED QUALIFICATIONS

\n

    \n
  • Experience designing clusters for large enterprise clients or neoclouds, not just internal infrastructure.
  • \n

  • Familiarity with NVIDIA NCP program requirements and certification processes.
  • \n

  • Experience with capacity or sparing modeling tools.
  • \n

Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...