Software Development Engineer Ai/Ml, Inference Serving, Aws Neuron

22 ore fa


Montà, Italia Amazon A tempo pieno

Software Development Engineer AI/ML, Inference Serving, AWS NeuronAWS Neuron is the software stack powering AWS Inferentia and Trainium machine learning accelerators, designed to deliver high-performance, low-cost inference at scale.The Neuron Serving team develops infrastructure to serve modern machine learning models—including large language models (LLMs) and multimodal workloads—reliably and efficiently on AWS silicon.We are seeking a Software Development Engineer to lead and architect our next-generation model serving infrastructure, with a particular focus on large-scale generative AI applications.Key job responsibilitiesArchitect and lead the design of distributed ML serving systems optimized for generative AI workloadsDrive technical excellence in performance optimization and system reliability across the Neuron ecosystemDesign and implement scalable solutions for both offline and online inference workloadsLead integration efforts with frameworks such as vLLM, SGLang, Torch XLA, TensorRT, and TritonDevelop and optimize system components for tensor/data parallelism and disaggregated servingImplement and optimize custom PyTorch operators and NKI kernelsMentor team members and provide technical leadership across multiple work streamsDrive architectural decisions that impact the entire Neuron serving stackCollaborate with customers, product owners, and engineering teams to define technical strategyAuthor technical documentation, design proposals, and architectural guidelinesA day in the lifeYou'll lead critical technical initiatives while mentoring team members.You'll collaborate with cross-functional teams of applied scientists, system engineers, and product managers to architect and deliver state-of-the-art inference capabilities.Your day might involve:Leading design reviews and architectural discussionsRapidly prototyping software to show customer valueDebugging complex performance issues across the stackMentoring junior engineers on system design and optimizationCollaborating with research teams on new ML serving capabilitiesDriving technical decisions that shape the future of Neuron's inference stackAbout the teamThe Neuron Serving team is at the forefront of scalable and resilient AI infrastructure at AWS.We focus on developing model-agnostic inference innovations, including disaggregated serving, distributed KV cache management, CPU offloading, and container-native solutions.Our team is dedicated to upstreaming Neuron SDK contributions to the open-source community, enhancing performance and scalability for AI workloads.We're committed to pushing the boundaries of what's possible in large-scale ML serving.Recent sharesQualifications5+ years of programming using a modern programming language such as Java, C++, or C#, including object-oriented design experience5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience5+ years of non-internship professional software development experienceExperience as a mentor, tech lead or leading an engineering teamPreferred QualificationsMaster's degree in computer science or equivalentDeep expertise in ML Frameworks/Libraries such as JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, TensorRT.Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies.Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position.These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation.Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Our inclusive culture empowers Amazonians to deliver the best results for our customers.If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information.If the country/region you're applying in isn't listed, please contact your Recruiting Partner.Our compensation reflects the cost of labor across several US geographic markets.The base pay for this position ranges from $151,300/year in our lowest geographic market up to $261,500/year in our highest geographic market.Pay is based on a number of factors including market location and may vary depending on job-related knowledge, skills, and experience.Amazon is a total compensation company.Dependent on the position offered, equity, sign-on payments, and other forms of compensation may be provided as part of a total compensation package, in addition to a full range of medical, financial, and/or other benefits.For more information, please visit .This position will remain posted until filled.Applicants should apply via our internal or external career site.Posted: December 11, **** (Updated 11 minutes ago)#J-*****-Ljbffr



  • Montà, Italia Amazon A tempo pieno

    Machine Learning Performance Engineer, Annapurna LabsOur team is responsible for the AWS Neuron software stack, which powers Generative AI and other advanced ML workloads on AWS's custom-built ML accelerators — Inferentia and Trainium. These accelerators deliver best-in-class performance and cost-efficiency for ML inference and training in the cloud.We're...


  • Montà, Italia Amazon A tempo pieno

    Our team is responsible for the AWS Neuron Compiler, a cutting-edge deep learning compiler stack that powers Generative AI and other advanced ML workloads on AWS's custom-built ML accelerators — Inferentia and Trainium.These accelerators deliver best-in-class performance and cost-efficiency for ML inference and training in the cloud.We're building a new...


  • Montà, Italia Canonical A tempo pieno

    Python and Kubernetes Software Engineer - Data, AI/ML & AnalyticsJoin to apply for the Python and Kubernetes Software Engineer - Data, AI/ML & Analytics role at CanonicalPython and Kubernetes Software Engineer - Data, AI/ML & AnalyticsJoin to apply for the Python and Kubernetes Software Engineer - Data, AI/ML & Analytics role at CanonicalCanonical is a...


  • Montà, Italia Amazon A tempo pieno

    Embedded Software Engineer, AWS Annapurna LabsAWS Utility Computing (UC) provides product innovations — from foundational services such as Amazon's Simple Storage Service (S3) and Amazon Elastic Compute Cloud (EC2), to consistently released new product innovations that continue to set AWS's services and features apart in the industry.As a member of the UC...


  • Montà, Italia Amazon A tempo pieno

    AWS Utility Computing (UC) provides product innovations— from foundational services such as Amazon's Simple Storage Service (S3) and Amazon Elastic Compute Cloud (EC2), to consistently released new product innovations that continue to set AWS's services and features apart in the industry.As a member of the UC organization, you'll support the development...

  • Machine Learning Engineer

    1 settimana fa


    Montà, Italia Synergen Health A tempo pieno

    Join to apply for theMachine Learning Engineerrole atSYNERGEN Health.SYNERGEN Health is a US-based company that pioneers comprehensive financial solutions for healthcare organizations, providing tech-enabled services, advanced analytics, FinTech payment solutions, machine learning, consulting, and software solutions.The company was ranked among the ****...

  • Software Engineer Leader

    1 settimana fa


    Montà, Italia Drivesec - We Secure Your Things A tempo pieno

    Role DescriptionIn this pivotal role, the Software Engineer Leader is responsible for driving the development of innovative software solutions that meet the needs of our clients and enhance our product offerings. The figure possesses a strong technical background in software development, combined with exceptional leadership skills to mentor and guide a team...


  • Montà, Italia Ion A tempo pieno

    Software Developer/Engineer - Graduate Development ProgramJoin to apply for the Software Developer/Engineer - Graduate Development Program role at IONSoftware Developer/Engineer - Graduate Development ProgramJoin to apply for the Software Developer/Engineer - Graduate Development Program role at IONGet AI-powered advice on this job and more exclusive...

  • Fea Ai

    22 ore fa


    Montà, Italia Innovative Numerics Llc A tempo pieno

    Job description:Job Title: CAE Machine Learning EngineerCompany: Innovative Numerics (US-based)Location: Remote Role SummaryWe are seeking a CAE Machine Learning Engineer to help accelerate Innovative Numerics' product roadmap by designing, building, and validating capabilities for structural simulation and mechanical design workflows.This role spans applied...


  • Montà, Italia Segula Technologies A tempo pieno

    Company DescriptionJoin the world of the future in a fast growing international company!At SEGULA Technologies you will have the opportunity to work on exciting projects and help shaping the future within an engineering company which is at the heart of innovation.From 3D printing, augmented reality, connected vehicle to the factory of the future – new...