Senior Data Engineer

2 ore fa

Albenga, Liguria, Italia MSD Tempo pieno

POSITION SUMMARY:

This position is responsible for managing and organizing data to support all business processes to achieve corporate and departmental goals. This position is responsible for identifying trends and communicating these trends clearly to others in the organization to ensure data is properly used. Core activities include troubleshooting data issues, and assisting the data architect to develop, align, and maintain architectures with business requirements. In addition, this position is expected to implement strategies to acquire high quality data, timely, accurately, reliably, and efficiently by developing the right data processes with the applications and integration engineers to meet all necessary compliance and governance needs.

DUTIES AND RESPONSIBILITIES:

  • Develop data pipelines from various data sources to target locations — including MES, ERP, PLM, SCADA/historian, and quality systems — to support the enterprise Digital Thread, including formatting, cleaning, and updating data as the business needs.
  • Develop and maintain accurate data structure and mapping documentation, including metadata aligned to specific business requirements.
  • Design and publish domain-oriented, reusable data products aligned to a data mesh model with clear ownership, SLAs, and discoverability — standardizing data structure and types across ETL/ELT processes.
  • Understand the big picture, set the right scope, and collaborate with the data architect, application engineers, integration specialists, business analysts, and other technical experts to ensure adequate content delivery to authorized users in a timely, effective, and secure manner.
  • Manage the full life cycle development for the current ETL/ELT deployments, applying software-engineering practices to data pipelines — version control (Git), automated testing, code review, CI/CD, and Infrastructure-as-Code (e.g., Terraform, CloudFormation) — to ensure repeatable, auditable deployments across environments.
  • Partner with manufacturing operations, quality, and supply chain teams to deliver analytics on OEE, yield, scrap, throughput, engineering-change cycle time, time-to-release, and end-to-end product genealogy/traceability supported by the Digital Thread.
  • Prepare AI/ML-ready datasets — including feature pipelines, dataset versioning and curated knowledge sources for predictive quality, anomaly detection, and retrieval-augmented (RAG) use cases — in partnership with data science and AI teams.

EXPERIENCE AND QUALIFICATIONS:

  • Bachelor’s degree in Computer Science, Software Engineering, or other related science or engineering discipline is required. Advanced degree preferred.
  • A minimum of four years experience in Data Modeling, Data Solution Development, and Data Integration.
  • Experience integrating data from manufacturing and operational systems (MES, ERP, PLM, SCADA/historian, LIMS, or QMS) is highly preferred.
  • Extensive experience working with data science tools/technologies, particularly Python, SQL, and/or C# .NET.
  • Experience analyzing business requirements, planning, executing actions, and solving complex problems.
  • Hands-on experience with the AWS data stack (S3, Glue, EMR, Redshift, Lake Formation, Kinesis, Lambda, and IAM) for building, securing, and operating production data platforms is required; AWS Certified Data Engineer or AWS Certified Data Analytics certification is highly preferred.
  • Experience contributing to a data mesh, data fabric, or Digital Thread initiative — including building domain-oriented data products, using a data catalog, and applying federated governance — is highly preferred.
  • Experience implementing data governance, lineage, and cataloging on AWS (e.g., AWS Glue Data Catalog, Lake Formation, and tools such as Collibra, Alation, Atlan, or OpenMetadata) is highly preferred.

KNOWLEDGE, SKILLS AND ABILITIES:

  • Demonstrated ability to manipulate and aggregate structured and unstructured data from multiple sources by constructing complex queries for analysis and reporting.
  • Excellent communication and interpersonal skills, with the ability to convey technical issues, tradeoffs, and results to technical and business stakeholders.
  • Strong ownership and delivery orientation — proactive, detail-oriented, and able to manage multiple priorities against time-sensitive deadlines.
  • Knowledge of ETL/ELT process tools, such as SSIS, Informatica, Talend, dbt, Fivetran, and/or Airflow is highly preferred.
  • Working knowledge of DataOps practices and tooling — including Git-based workflows, CI/CD (e.g., GitHub Actions, GitLab CI, AWS CodePipeline), Infrastructure-as-Code (Terraform or CloudFormation), automated data testing, and pipeline observability — is highly preferred.
  • Working knowledge of modern data platform technologies — such as cloud