Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training
2 giorni fa
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training Annapurna Labs designs silicon and software that accelerate innovation. Customers choose us to create cloud solutions that solve challenges that were unimaginable a short time ago—even yesterday. Our custom chips, accelerators, and software stacks enable us to take on technical challenges that have never been seen before, and deliver results that help our customers change the world. AWS Neuron is the complete software stack for the AWS Trainium (Trn1/Trn2) and Inferentia (Inf1/Inf2) cloud‑scale Machine Learning accelerators. This role is for a Senior Machine Learning Engineer in the Distributed Training team for AWS Neuron, responsible for development, enablement and performance tuning of a wide variety of ML model families, including massive‑scale Large Language Models (LLM) such as GPT and Llama, as well as Stable Diffusion, Vision Transformers (ViT) and many more. Key Responsibilities Lead efforts to build distributed training support into PyTorch and JAX using XLA, the Neuron compiler, and runtime stacks. Optimize models to achieve peak performance and maximize efficiency on AWS custom silicon, including Trainium and Inferentia, as well as Trn2, Trn1, Inf1, and Inf2 servers. Work with cross‑functional teams (chip architects, compiler engineers, runtime engineers) to create, build and tune distributed training solutions with Trainium instances. Deep dive into model training pipelines and extend distributed training libraries such as FSDP, Deepspeed, Nemo, and others for the Neuron system. About the Team Our team collaborates closely with hardware and software experts to support new members, promote knowledge sharing, and ensure high‑quality code through mentoring and thorough code reviews. We value diversity and inclusive culture, providing flexibility, mentorship, and growth opportunities for all team members. Basic Qualifications Bachelor’s degree in computer science or equivalent. 5+ years of noninternship professional software development experience. 5+ years of programming in at least one software programming language. 5+ years of leading design or architecture of new and existing systems. 5+ years of full software development life cycle experience, including coding standards, code reviews, source control management, build processes, testing, and operations. Experience as a mentor, tech lead or engineering team lead. Experience in machine learning, data mining, information retrieval, statistics or natural language processing. Preferred Qualifications Master’s degree in computer science or equivalent. Experience in computer architecture. Previous software engineering expertise with PyTorch/JAX/TensorFlow, distributed libraries and frameworks, and end‑to‑end model training. Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Location: US, Massachusetts, North Reading. Salary: $193,300 - $261,500 annually. Benefits include health insurance, 401(k) matching, paid time off, parental leave, and more. View full benefits at #J-18808-Ljbffr
-
torino, Italia Amazon A tempo pienoSr. Software Engineer- AI/ML, AWS Neuron Distributed Training Annapurna Labs designs silicon and software that accelerate innovation. Customers choose us to create cloud solutions that solve challenges that were unimaginable a short time ago—even yesterday. Our custom chips, accelerators, and software stacks enable us to take on technical challenges that...
-
ML Systems Engineer: Distributed Training on Neuron/Trainium
2 settimane fa
Torino, Italia Amazon A tempo pienoA leading technology company in Torino is seeking a Software Engineer to join the Machine Learning Applications team for AWS Neuron. This role will focus on developing and tuning distributed training solutions for large models, including PyTorch and JAX support. The ideal candidate will have extensive experience in software development, machine learning, and...
-
Software Development Engineer
4 giorni fa
Torino, Italia Amazon A tempo pienoSoftware Development Engineer – AI/ML, AWS Neuron, Multimodal Inference The Annapurna Labs team at Amazon Web Services (AWS) builds AWS Neuron, the software development kit used to accelerate deep learning and GenAI workloads on Amazon's custom machine learning accelerators, Inferentia and Trainium. The AWS Neuron SDK is the backbone for accelerating deep...
-
Software Development Engineer
5 giorni fa
Torino, Italia Amazon A tempo pienoSoftware Development Engineer – AI/ML, AWS Neuron, Multimodal Inference The Annapurna Labs team at Amazon Web Services (AWS) builds AWS Neuron, the software development kit used to accelerate deep learning and GenAI workloads on Amazon’s custom machine learning accelerators, Inferentia and Trainium. The AWS Neuron SDK is the backbone for accelerating...
-
Software Development Engineer
3 giorni fa
Sant'Ambrogio di Torino, Italia Amazon A tempo pienoSoftware Development Engineer – AI/ML, AWS Neuron, Multimodal InferenceThe Annapurna Labs team at Amazon Web Services (AWS) builds AWS Neuron, the software development kit used to accelerate deep learning and GenAI workloads on Amazon’s custom machine learning accelerators, Inferentia and Trainium. The AWS Neuron SDK is the backbone for accelerating deep...
-
Software Engineer- AI/ML, AWS Neuron
1 settimana fa
torino, Italia Amazon A tempo pienoAWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud‑scale machine learning accelerators and the Trn1 and Inf1 servers that use them. This role is for a software engineer in the Machine Learning Applications (ML Apps) team for AWS Neuron. This role is responsible for development, enablement and performance tuning of a wide...
-
Software Engineer- AI/ML, AWS Neuron
1 settimana fa
torino, Italia Amazon A tempo pienoAWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud‑scale machine learning accelerators and the Trn1 and Inf1 servers that use them. This role is for a software engineer in the Machine Learning Applications (ML Apps) team for AWS Neuron. This role is responsible for development, enablement and performance tuning of a wide...
-
Software Engineer- AI/ML, AWS Neuron
2 settimane fa
Torino, Italia Amazon A tempo pienoAWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud‑scale machine learning accelerators and the Trn1 and Inf1 servers that use them. This role is for a software engineer in the Machine Learning Applications (ML Apps) team for AWS Neuron. This role is responsible for development, enablement and performance tuning of a wide...
-
AI/ML Software Engineer
4 giorni fa
Torino, Italia Amazon A tempo pienoA global technology company in Italy is seeking a Software Development Engineer specializing in AI/ML for AWS Neuron. The ideal candidate should have over 3 years of software development experience, specifically in C++ and Python, and a solid grasp of machine learning fundamentals. Responsibilities include optimizing machine-learning models for AWS's custom...
-
Torino, Italia Amazon A tempo pienoA global technology company in Italy is seeking a Software Development Engineer specializing in AI/ML for AWS Neuron. The ideal candidate should have over 3 years of software development experience, specifically in C++ and Python, and a solid grasp of machine learning fundamentals. Responsibilities include optimizing machine-learning models for AWS's custom...