Senior Infrastructure Engineer — OpenStack

1 giorno fa

Palermo, Sicily, Italia Azienda Anonima Tempo pieno

SENIOR INFRASTRUCTURE ENGINEER - OPENSTACK & CLOUD SYSTEMS

We are looking for a Senior Infrastructure Engineer to serve as a technical leader within our Infrastructure Team, driving the design, evolution, and operation of the software-defined infrastructure that powers some of the largest HPC and AI clusters in Europe. OpenStack is the primary production platform today; the role owns the broader Linux, Kubernetes, and bare-metal systems stack around it.

We hire first for depth of systems engineering — the ability to understand why something does not work, not just how to restart it: reading an strace, reading the source of a service, isolating a kernel or network issue, and driving a fix upstream when needed. A strong engineer with this foundation who knows OpenStack — or an equivalent large-scale production platform — is exactly who we are looking for; the specific stack can be learned, the way of reasoning cannot.

The successful candidate will be responsible for the architecture and operations of cloud control planes in large-scale production environments, ensuring reliability, scalability, and operational continuity. They will lead critical technical initiatives, plan and execute infrastructure upgrades and migrations, collaborate with leading technology vendors to manage escalations, and provide technical leadership to the team through mentoring, knowledge sharing, and best practices.

KEY RESPONSIBILITIES:

  • Own the architecture and day-2 operations of OpenStack control planes for clusters of hundreds of nodes today, with a growth path to 1000+ nodes at upcoming public and private AI Factories: uptime, capacity, performance, and security KPIs.
  • Diagnose and resolve complex, cross-layer production issues down to the root cause — kernel, systemd, storage, and network stack — using tools such as strace, perf, and packet capture, reading service source code where needed and contributing fixes upstream.
  • Design and operate the advanced networking underpinning high-throughput HPC and AI workloads (Neutron OVN / OVS, SR-IOV, VF-LAG, DPDK, BGP-EVPN), and drive technology and architecture decisions with senior team members and stakeholders.
  • Coordinate deployments and upgrades across geographically distributed sites, including cross-site data replication, federated identity, and disaster-recovery posture.
  • Develop and maintain advanced Infrastructure-as-Code pipelines (Ansible, OpenTofu / Terraform, Helm) and enforce gitops-style review for production change management.
  • Produce and maintain technical documentation, including operational runbooks — step-by-step procedures with explicit go / no-go decision gates and rollback plans — for control-plane upgrades, security patching, and migrations, as well as RCA reports.
  • Mentor mid-level and junior engineers on Linux and OpenStack internals, lifecycle operations, and production best practices; collaborate with the Presales team in designing systems end-to-end from hardware configuration to software stack.

EDUCATION:

Master's degree or Ph.D. in Computer Science, Telecommunications Engineering, Network Engineering, or a related STEM field — or equivalent practical experience, including:

  • 5+ years of hands-on engineering experience in Linux systems, cloud architecture, network engineering, and complex production systems.
  • 3+ years of OpenStack production experience at scale (multi-site or strict SLA) — or equivalent experience operating a large-scale production platform, with the depth to become productive on OpenStack within the first months.
  • 2+ years of direct vendor escalation accountability in enterprise or hyperscaler environments.

SKILLS AND COMPETENCES - CORE TECHNICAL SKILLS

Linux & systems engineering (deep):

  • Expert Linux sysadmin (Rocky / RHEL, Ubuntu, SLES);
  • Advanced troubleshooting strace, perf, ftrace / eBPF, gdb; comfortable reading the source of the services operated and contributing fixes upstream;
  • Container platforms (Docker, Podman, Singularity) and their runtime internals.

OpenStack & cloud virtualization

  • Strong, hands-on proficiency on Neutron, Nova, Ironic, Cinder, Keystone, Manila.
  • Direct hands-on experience with Kayobe + Kolla-Ansible is a strong plus.

Advanced networking

  • Working knowledge of BGP / OSPF / EVPN / VXLAN, Spine-Leaf datacenter fabrics, SDN, OVN / OVS.
  • SR-IOV + VF-LAG, hardware offload, DPDK; Mellanox / low-latency networking.

Infrastructure as Code & Automation

  • Production-quality automation in Bash, Python, or Go.
  • Hands-on Ansible (modules / roles / collections); Terraform / OpenTofu modules; Helm.

Container orchestration (a plus)

  • Production-grade Kubernetes (bare metal and over OpenStack), with focus on GPU-accelerated workloads; Cluster API, ArgoCD.