Senior Cloud Infrastructure and Network Operations Solutions Architect

6 ore fa

roma, lazio, Italia Nvidia Tempo pieno
Overview

In this role you will own Day 2 fabric operations for NVIDIA’s Cloud Partner estates, guiding the reliability and performance of InfiniBand, RoCE/Ethernet, and NVLink/NVSwitch fabrics at scale. You’ll work with cross-functional teams and customers to architect, validate, and operate complex GPU-centric networks across multi-tenant environments. This is a hands-on leadership role in a fast-moving, open-source–first fabric ecosystem. You’ll drive operational excellence, incident response, and knowledge transfer to partner teams, shaping how large AI/HPC systems run reliably.

Responsabilità
  • Own Day 2 fabric operations across NVIDIA Cloud Partner fleets (NVLink/NVSwitch partition management, maintenance-partition isolation, safe partition-change workflows)
  • Manage switch software and firmware lifecycle including upgrades and rollout campaigns with minimal production disruption
  • Support Day 1 fabric validation and acceptance (InfiniBand/UFM bring-up, Spectrum-X/RoCE config, burn-in against MTBI/goodput targets)
  • Minimise handover time to production by reducing duplicated validation across hardware bring-up and partner operations
  • Drive fleet-wide fabric reliability through telemetry, fault detection, remediation, and root-cause analysis of network-induced job failures
  • Provide consultative guidance and hands-on troubleshooting across fabric stack (NICs/DPUs, switch OS, routing, Kubernetes networking) and lead knowledge transfer with runbooks for partner teams
  • Serve as technical leader for assigned accounts and present to executive stakeholders
Requisiti fondamentali
  • 5+ years in data center networking, fabric engineering, or large-scale network operations
  • Deep understanding of data center architectures and RDMA fabrics (InfiniBand, RoCE/Ethernet) including topology, routing, congestion control, and QoS
  • Hands-on experience with InfiniBand/UFM, Spectrum-X Ethernet, Cumulus Linux and/or SONiC, ConnectX/BlueField NICs/DPUs, NVLink/NVSwitch on NVL72-class platforms
  • Strong Linux knowledge (RedHat/Ubuntu), switch OS internals, security, and HPC/AI traffic patterns
  • Automation and Observability skills: Python/Bash, IaC tools (Ansible, Terraform), GitOps, Grafana/Loki/Prometheus
  • Proven ability to measure and improve MTBI and job goodput, and to lead architectural reviews with executive stakeholders
  • Experience with Kubernetes networking in GPU clusters and multi-tenant network isolation concepts
  • Strong consultative mindset with demonstrated capacity to transfer knowledge and create runbooks for partner teams
  • Consultative mindset
  • Strong communication and presentation skills
  • Cross-functional collaboration
  • InfiniBand and UFM
  • Spectrum-X Ethernet
  • Cumulus Linux, SONiC