Senior Cloud Infrastructure and Network Operations Solutions Architect
Salva questo lavoro e mantieni la tua ricerca organizzata
Crea un account gratuito per salvare lavori, creare avvisi e tornare a questa inserzione dalla tua dashboard.
Continuando accetti i nostri Termini & Informativa sulla privacy.
In this role you will own Day 2 fabric operations for NVIDIA’s Cloud Partner estates, guiding the reliability and performance of InfiniBand, RoCE/Ethernet, and NVLink/NVSwitch fabrics at scale. You’ll work with cross-functional teams and customers to architect, validate, and operate complex GPU-centric networks across multi-tenant environments. This is a hands-on leadership role in a fast-moving, open-source–first fabric ecosystem. You’ll drive operational excellence, incident response, and knowledge transfer to partner teams, shaping how large AI/HPC systems run reliably.
Responsabilità- Own Day 2 fabric operations across NVIDIA Cloud Partner fleets (NVLink/NVSwitch partition management, maintenance-partition isolation, safe partition-change workflows)
- Manage switch software and firmware lifecycle including upgrades and rollout campaigns with minimal production disruption
- Support Day 1 fabric validation and acceptance (InfiniBand/UFM bring-up, Spectrum-X/RoCE config, burn-in against MTBI/goodput targets)
- Minimise handover time to production by reducing duplicated validation across hardware bring-up and partner operations
- Drive fleet-wide fabric reliability through telemetry, fault detection, remediation, and root-cause analysis of network-induced job failures
- Provide consultative guidance and hands-on troubleshooting across fabric stack (NICs/DPUs, switch OS, routing, Kubernetes networking) and lead knowledge transfer with runbooks for partner teams
- Serve as technical leader for assigned accounts and present to executive stakeholders
- 5+ years in data center networking, fabric engineering, or large-scale network operations
- Deep understanding of data center architectures and RDMA fabrics (InfiniBand, RoCE/Ethernet) including topology, routing, congestion control, and QoS
- Hands-on experience with InfiniBand/UFM, Spectrum-X Ethernet, Cumulus Linux and/or SONiC, ConnectX/BlueField NICs/DPUs, NVLink/NVSwitch on NVL72-class platforms
- Strong Linux knowledge (RedHat/Ubuntu), switch OS internals, security, and HPC/AI traffic patterns
- Automation and Observability skills: Python/Bash, IaC tools (Ansible, Terraform), GitOps, Grafana/Loki/Prometheus
- Proven ability to measure and improve MTBI and job goodput, and to lead architectural reviews with executive stakeholders
- Experience with Kubernetes networking in GPU clusters and multi-tenant network isolation concepts
- Strong consultative mindset with demonstrated capacity to transfer knowledge and create runbooks for partner teams
- Consultative mindset
- Strong communication and presentation skills
- Cross-functional collaboration
- InfiniBand and UFM
- Spectrum-X Ethernet
- Cumulus Linux, SONiC