AI Platform Engineer
1 giorno fa
Turin, Piedmont, Italia
Reply
Tempo pieno
50.000 € - 75.000 € Contratto
Gratuito con email o Google
Salva questo lavoro e mantieni la tua ricerca organizzata
Crea un account gratuito per salvare lavori, creare avvisi e tornare a questa inserzione dalla tua dashboard.
Gratuito con email o Google
Continuando accetti i nostri Termini & Informativa sulla privacy.
Are you an AI Platform Engineer expert in the platform that turns a researcher's idea into a running job on a cluster?
Join Reply
WHO WE ARE
Reply is a network of highly specialized companies which support leading industrial groups in defining and developing business models using new technology and communication paradigms, such as Big Data, Cloud Computing, Digital Communication, Internet of Things, Mobile and Social Networking. Reply focuses on Consultancy, System Integration and Application Management, covering three areas of competence: Processes, Applications and Technologies.
WHAT YOU WILL FIND
• Core activities. You will keep the core of the Model Factory platform standing, and you will make it grow. You will orchestrate GPU workloads and distributed training that hold up under load, build versioning and governance across data, models and evaluation runs so results are still reproducible months later, and shorten the distance between a researcher's idea and a running job on the cluster. Every model we train, distil or optimise — for our own portfolio and for client engagements — runs on what you build.
• Tech & tools stack. You will work SkyPilot, Dagster, NVIDIA NeMo, Kubernetes, MLflow and Unity Catalog, the core of the platform. Around them: the NVIDIA stack (CUDA, NCCL), job scheduling across national (Italian) and European HPC clusters and cloud (such as Scaleway or Nebius), infrastructure as code and GitOps CI/CD, vLLM for inference serving, and Prometheus and Grafana observability that tells us what a run cost and why it failed.
• Mentoring, collaboration, and continuous growth. You will work alongside the AI research engineers who train the models, the model design engineers who specify them, the project managers who commit the dates, and cloud and security specialists — plus client platform teams when the environment lands on their side. Your first customers are internal: if the platform is slow, opaque or fragile, they feel it long before any client does.
WHAT WE OFFER
• An offer tailored to your experience. This position is open to people with varying levels of expertise and seniority. The compensation package will be determined based on your professional background, technical skills, expertise, and the level of responsibility associated with the role. The collective labour agreement (CCNL) applied is the Italian Metalworking Industry Agreement. The job classification will be assessed starting from level B2, and the gross annual salary (RAL) will range from €50,000 to €75,000. Previous experience with cutting-edge technologies such as GPU cluster orchestration, distributed training on national and European HPC infrastructure, SkyPilot, Dagster and NVIDIA NeMo, LLM inference optimisation, and data, model and evaluation governance in MLflow and Unity Catalog will be considered a strong asset.
• A structured career path. Our Career Path offers opportunities to grow both as a technical specialist and as a future leader. Your ambitions, the skills you develop, and the results you achieve will shape your professional journey.
• Continuous learning. Technology evolves rapidly
- and so do we. You will join an environment that encourages curiosity, continuous learning, and the exploration of new ideas and emerging technologies.
• The benefits of being a Replyer. You will also have access to the benefits and initiatives dedicated to our community. WHAT WE NEED
• Academic background. Degree in Computer Engineering, Computer Science or a related technical field. Solid fundamentals in distributed systems, networking matter more to us than any specific coursework.
• Technical & strategic skills. 3-8 years of professional experience. Kubernetes and distributed training that genuinely work in your hands: you can debug a training job that hangs, size a GPU allocation, and tell a scheduling problem from a networking one. Strong Python, and infrastructure as code as a default rather than an afterthought. Having run ML platforms at scale is required — we care that the fundamentals are real and that you design for reproducibility, multi-tenancy and cost attribution from day one. The position is open to varying levels of expertise and seniority.
• Soft skills. You should be pragmatic and ownership-driven, comfortable being the person everyone else depends on. Clear written communication, a bias for automation over heroics, and the patience to make someone else's workflow actually work.
• Nice to have. Experience with SkyPilot, Dagster or NVIDIA NeMo in production; HPC schedulers (Slurm) and InfiniBand/RDMA networking; LLM inference optimisation (quantisation, continuous batching, KV-cache); Unity Catalog or Databricks governance; open-source contributions to the ML infrastructure ecosystem; sovereign or on-premise AI environments. WHAT ARE THE NEXT STEPS The first step of our recruiting process will be the meetings with the technical referents and then a face to face interview with the HR team. We care about an equal recruiting process. Feel interested? Reply is committed to embracing diversity and creating an inclusive work environment by valuing the uniqueness of people regardless of age, gender, sexual orientation, religion, nationality, or disabilities as protected by Italian Law (L.68/99). Furthermore, Reply is committed to ensuring a fair and accessible selection process: to help you during the recruitment process, please let us know of any kind of support you may need.
• Core activities. You will keep the core of the Model Factory platform standing, and you will make it grow. You will orchestrate GPU workloads and distributed training that hold up under load, build versioning and governance across data, models and evaluation runs so results are still reproducible months later, and shorten the distance between a researcher's idea and a running job on the cluster. Every model we train, distil or optimise — for our own portfolio and for client engagements — runs on what you build.
• Tech & tools stack. You will work SkyPilot, Dagster, NVIDIA NeMo, Kubernetes, MLflow and Unity Catalog, the core of the platform. Around them: the NVIDIA stack (CUDA, NCCL), job scheduling across national (Italian) and European HPC clusters and cloud (such as Scaleway or Nebius), infrastructure as code and GitOps CI/CD, vLLM for inference serving, and Prometheus and Grafana observability that tells us what a run cost and why it failed.
• Mentoring, collaboration, and continuous growth. You will work alongside the AI research engineers who train the models, the model design engineers who specify them, the project managers who commit the dates, and cloud and security specialists — plus client platform teams when the environment lands on their side. Your first customers are internal: if the platform is slow, opaque or fragile, they feel it long before any client does.
WHAT WE OFFER
• An offer tailored to your experience. This position is open to people with varying levels of expertise and seniority. The compensation package will be determined based on your professional background, technical skills, expertise, and the level of responsibility associated with the role. The collective labour agreement (CCNL) applied is the Italian Metalworking Industry Agreement. The job classification will be assessed starting from level B2, and the gross annual salary (RAL) will range from €50,000 to €75,000. Previous experience with cutting-edge technologies such as GPU cluster orchestration, distributed training on national and European HPC infrastructure, SkyPilot, Dagster and NVIDIA NeMo, LLM inference optimisation, and data, model and evaluation governance in MLflow and Unity Catalog will be considered a strong asset.
• A structured career path. Our Career Path offers opportunities to grow both as a technical specialist and as a future leader. Your ambitions, the skills you develop, and the results you achieve will shape your professional journey.
• Continuous learning. Technology evolves rapidly
- and so do we. You will join an environment that encourages curiosity, continuous learning, and the exploration of new ideas and emerging technologies.
• The benefits of being a Replyer. You will also have access to the benefits and initiatives dedicated to our community. WHAT WE NEED
• Academic background. Degree in Computer Engineering, Computer Science or a related technical field. Solid fundamentals in distributed systems, networking matter more to us than any specific coursework.
• Technical & strategic skills. 3-8 years of professional experience. Kubernetes and distributed training that genuinely work in your hands: you can debug a training job that hangs, size a GPU allocation, and tell a scheduling problem from a networking one. Strong Python, and infrastructure as code as a default rather than an afterthought. Having run ML platforms at scale is required — we care that the fundamentals are real and that you design for reproducibility, multi-tenancy and cost attribution from day one. The position is open to varying levels of expertise and seniority.
• Soft skills. You should be pragmatic and ownership-driven, comfortable being the person everyone else depends on. Clear written communication, a bias for automation over heroics, and the patience to make someone else's workflow actually work.
• Nice to have. Experience with SkyPilot, Dagster or NVIDIA NeMo in production; HPC schedulers (Slurm) and InfiniBand/RDMA networking; LLM inference optimisation (quantisation, continuous batching, KV-cache); Unity Catalog or Databricks governance; open-source contributions to the ML infrastructure ecosystem; sovereign or on-premise AI environments. WHAT ARE THE NEXT STEPS The first step of our recruiting process will be the meetings with the technical referents and then a face to face interview with the HR team. We care about an equal recruiting process. Feel interested? Reply is committed to embracing diversity and creating an inclusive work environment by valuing the uniqueness of people regardless of age, gender, sexual orientation, religion, nationality, or disabilities as protected by Italian Law (L.68/99). Furthermore, Reply is committed to ensuring a fair and accessible selection process: to help you during the recruitment process, please let us know of any kind of support you may need.