Data Engineer Principal_3003

6 giorni fa

rome, lazio, Italia Allianz Tempo pieno
Data Engineer Principal_3003 Salary: €110,000 - 150,000 yearly
Azienda: Allianz
Tipo Lavoro: Full Time Italy
Descrizione Lavoro - Data Engineer Principal_3003 We are looking for an experienced Principal Data Engineer to work hands-on on the re-engineering of an existing enterprise data platform built on Azure Synapse Analytics. The role requires strong technical depth to audit, understand, and validate a complex end-to-end data architecture spanning source ingestion through to consumption — and to help deliver the migration of validated workloads to Databricks. This is hands on role and demands the ability to reverse-engineer existing implementations, assess their correctness, and build to an agreed migration strategy.
Key Responsibilities Contribute to the technical assessment and re-engineering of an existing enterprise data platform, spanning all layers from source ingestion through to data consumption
Reverse-engineer, document, and validate existing pipeline logic, data models, transformation frameworks, and data governance controls
Identify gaps, defects, and technical debt across the platform and remediate where implementations are incorrect or sub-optimal
Ensure correctness of data processing patterns including change data capture, slowly changing dimensions, deduplication, and business reconciliation
Implement target-state designs aligned to modern lakehouse principles, ensuring feature parity and business logic fidelity during transitions
Support platform evolution initiatives, including parallel-run phases where multiple implementations operate simultaneously, validating output consistency before cutover
Execute migration of existing workloads to modern data platforms, preserving existing governance and control framework semantics
Re-implement ingestion, transformation, and orchestration pipelines on target platforms, maintaining audit, quality, and reconciliation standards
Collaborate with business, data governance, and architecture stakeholders to validate embedded business rules and data quality requirements
Mentor less experienced engineers, review code and designs, and support decommission planning for legacy components
Core Technical Skills Azure Synapse & Data Platform Mandatory hands-on expertise with:
Azure Synapse Analytics (Pipelines, Spark Pool, Dedicated SQL Pool)
Azure Data Lake Storage Gen2 (ADLS Gen2)
Delta Lake on Azure (Synapse Lakehouse patterns)
Oracle Golden Gate Replication for real-time source integration
Azure Analysis Services and Power BI consumption layer patterns
Deep understanding of medallion architecture: Raw / Harmonized / Conformed / Consumption layers
Strong knowledge of SCD Type 0/1/2, CDC patterns, soft/hard delete, and retroactive change processing
Experience with Synapse SQL Pool — stored procedures, control tables, and data quality validation patterns
Experience with audit, balance, and control frameworks — parameterized, modular pipeline governance at enterprise scale
Familiarity with config-driven and automation-first pipeline patterns (YAML, PySpark, SQL-driven generation from mapping documents)
Databricks & Lakehouse Hands-on experience with Azure Databricks (Delta Live Tables, Unity Catalog preferred)
Strong Apache Spark skills (PySpark / Spark SQL)
Experience migrating workloads from legacy data warehouse or Synapse environments to a Databricks Lakehouse
Ability to re-implement governance and control frameworks natively in Databricks (audit logging, reconciliation, DQ checks)
Experience with Delta Lake features: MERGE, CDC, time travel, schema enforcement
Data Engineering & Development Strong Python and SQL programming skills
Experience with ETL/ELT at scale: denormalization, surrogate keys, directory tables, curated data models
Experience integrating complex data sources: Oracle DB, SQL Server, Azure SQL DB, file systems, Salesforce, APIs
Strong data modelling skills: relational, dimensional, and lakehouse-oriented
DevOps & Automation CI/CD pipelines for data engineering (Azure DevOps / GitHub Actions)
Infrastructure as Code (Terraform or ARM)
Containerization (Docker)
Experience with automated testing frameworks for data pipelines (unit testing, reconciliation-based validation)
Nice to Have Experience with Unity Catalog for data governance and lineage
Familiarity with Azure Purview for data cataloguing and governance
Exposure to real-time and streaming pipelines (Event Hub / Kafka / Kinesis)
Experience with GenAI or ML platform integration (MLOps, feature engineering pipelines)
Familiarity with monitoring and observability tools (e.g., Dynatrace)
Exposure to BI tools (Power BI, Tableau)
Experience & Profile 7+ years of hands-on experience in Data Engineering, including platform migration or re-engineering work
Proven track record working on existing, complex enterprise data platforms —