Senior Data Engineer
4 dage siden
Copenhagen, Capital Region of Denmark
Altimetrik
Fuldtid
Gratis med e-mail eller Google
Gem dette job og hold din søgning organiseret
Opret en gratis konto for at gemme job, oprette underretninger og vende tilbage til denne fortegnelse fra dit dashboard.
Gratis med e-mail eller Google
Altimetrik Poland is a digital enablement company. We deliver bite-size outcomes to enterprises and start-ups from all industries in an agile way to help them scale and accelerate their businesses. We are unique in Poland's IT market. Our differentiators are an innovation-first approach, a strong focus on core development, and an ability to attack the challenging and complex problems of the biggest companies in the world.
We are looking for an accomplished Senior Data Engineer with 10+ years of experience spanning the full modern data stack — Databricks Lakehouse, Snowflake Cloud Data Platform, AWS data services, and enterprise Data Quality & Observability. You will architect, build, and operate scalable, reliable data platforms that power analytics, ML, and business intelligence across the organization. This is a high-impact individual contributor and technical leadership role.
Key Responsibilities
Databricks & Lakehouse ● Design and build scalable data pipelines using Databricks and Apache Spark (PySpark/Scala) for batch and streaming workloads. ● Architect Lakehouse solutions with Delta Lake (ACID transactions, schema evolution, time travel) and medallion architecture (bronze/silver/gold). ● Build and maintain ETL/ELT workflows using Databricks Workflows, Delta Live Tables (DLT), and Auto Loader. ● Implement Unity Catalog for enterprise data governance, access control, and data lineage tracking. ● Collaborate with ML/Data Science teams on feature engineering pipelines, MLflow integration, and model deployment. Snowflake ● Architect Snowflake environments including virtual warehouses, databases, schemas, and storage integrations. ● Design and implement data pipelines using Snowflake Streams, Tasks, Snowpipe, Snowpark, and Dynamic Tables. ● Build data models (star/snowflake schemas, Data Vault 2.0) optimized for Snowflake performance. ● Implement Snowflake RBAC, Row-Level Security, Column Masking, data sharing, and cross-region replication. ● Optimize query performance through clustering keys, materialized views, result caching, and search optimization. Data Quality & Observability ● Design and implement enterprise-grade data quality frameworks covering profiling, validation, monitoring, and alerting. ● Define and enforce DQ rules — completeness, accuracy, consistency, timeliness, and uniqueness — across all pipelines. ● Build automated DQ checks using Great Expectations, Soda Core, Monte Carlo, or Deequ. ● Develop data observability pipelines to detect anomalies, schema drift, and volume/freshness issues. ● Instrument end-to-end data lineage (OpenLineage, Marquez) and integrate with data catalog platforms (Collibra, Atlan). ● Define data contracts, SLAs, and quality scorecards in collaboration with Data Governance teams. AWS Cloud Data Platform ● Architect and implement cloud-native data lakes and warehouses using S3, Glue, EMR, Redshift, Athena, and Lake Formation. ● Build real-time streaming pipelines using Amazon Kinesis, MSK (Kafka), and Lambda. ● Provision and manage data infrastructure using Terraform or AWS CDK (Infrastructure as Code). ● Implement fine-grained access control via Lake Formation, IAM, and Glue Data Catalog policies. ● Orchestrate data workflows using Apache Airflow (MWAA) or AWS Step Functions. Cross-Cutting ● Lead code reviews, establish engineering best practices, and mentor junior/mid-level data engineers. ● Integrate data pipelines with CI/CD tools (Git, GitHub Actions, Jenkins, CodePipeline) for automated deployment. ● Monitor and optimize cloud costs across Databricks, Snowflake, and AWS workloads. ● Translate business requirements into technical data platform solutions with clear documentation. ◆ Required Skills & Qualifications Core Engineering ● 10+ years of Data Engineering experience; expert-level hands-on across Databricks, Snowflake, and AWS. ● Proficiency in Python and SQL; strong command of PySpark and/or Scala for distributed data processing. ● Deep understanding of distributed computing, data lake architecture, and cloud-native design patterns. Databricks ● 4+ years with Databricks: Delta Lake, Delta Live Tables, Databricks Workflows, Unity Catalog, and Spark optimization. ● Experience with Databricks on AWS, Azure, or GCP — cluster tuning, auto-scaling, and cost controls. Snowflake ● 4+ years with Snowflake: virtual warehouses, micro-partitioning, Snowpark, Streams, Tasks, and Snowpipe. ● Hands-on with Snowflake data modeling (dimensional, Data Vault), RBAC, and compliance frameworks. ● Experience with ETL/ELT tools: dbt, Fivetran, Matillion, or Informatica integrated with Snowflake. Data Quality ● 3+ years building DQ frameworks using Great Expectations, Soda Core, Apache Griffin, or Deequ. ● Experience with observability platforms: Monte Carlo, Acceldata, or Atlan; knowledge of OpenLineage/Marquez. ● Understanding of data governance standards, GDPR/CCPA compliance, and regulatory requirements. AWS ● 5+ years on AWS data services: S3, Glue, EMR, Redshift, Athena, Kinesis, MSK, and Lake Formation. ● Hands-on with Terraform or AWS CDK; familiarity with AWS IAM, VPC, and security best practices. ● Experience with MWAA (Airflow), Step Functions, or equivalent orchestration tools. ◆ Good to Have ● Databricks Certified Data Engineer Professional or Associate. ● SnowPro Core or SnowPro Advanced Architect certification. ● AWS Certified Data Engineer – Associate or Solutions Architect – Professional. ● Experience with Apache Iceberg or Hudi on S3; knowledge of real-time lakehouse patterns. ● Exposure to ML pipeline orchestration: MLflow, SageMaker, or Vertex AI. ● Familiarity with data mesh architectures, federated governance, and data contracts.
Key Responsibilities
Databricks & Lakehouse ● Design and build scalable data pipelines using Databricks and Apache Spark (PySpark/Scala) for batch and streaming workloads. ● Architect Lakehouse solutions with Delta Lake (ACID transactions, schema evolution, time travel) and medallion architecture (bronze/silver/gold). ● Build and maintain ETL/ELT workflows using Databricks Workflows, Delta Live Tables (DLT), and Auto Loader. ● Implement Unity Catalog for enterprise data governance, access control, and data lineage tracking. ● Collaborate with ML/Data Science teams on feature engineering pipelines, MLflow integration, and model deployment. Snowflake ● Architect Snowflake environments including virtual warehouses, databases, schemas, and storage integrations. ● Design and implement data pipelines using Snowflake Streams, Tasks, Snowpipe, Snowpark, and Dynamic Tables. ● Build data models (star/snowflake schemas, Data Vault 2.0) optimized for Snowflake performance. ● Implement Snowflake RBAC, Row-Level Security, Column Masking, data sharing, and cross-region replication. ● Optimize query performance through clustering keys, materialized views, result caching, and search optimization. Data Quality & Observability ● Design and implement enterprise-grade data quality frameworks covering profiling, validation, monitoring, and alerting. ● Define and enforce DQ rules — completeness, accuracy, consistency, timeliness, and uniqueness — across all pipelines. ● Build automated DQ checks using Great Expectations, Soda Core, Monte Carlo, or Deequ. ● Develop data observability pipelines to detect anomalies, schema drift, and volume/freshness issues. ● Instrument end-to-end data lineage (OpenLineage, Marquez) and integrate with data catalog platforms (Collibra, Atlan). ● Define data contracts, SLAs, and quality scorecards in collaboration with Data Governance teams. AWS Cloud Data Platform ● Architect and implement cloud-native data lakes and warehouses using S3, Glue, EMR, Redshift, Athena, and Lake Formation. ● Build real-time streaming pipelines using Amazon Kinesis, MSK (Kafka), and Lambda. ● Provision and manage data infrastructure using Terraform or AWS CDK (Infrastructure as Code). ● Implement fine-grained access control via Lake Formation, IAM, and Glue Data Catalog policies. ● Orchestrate data workflows using Apache Airflow (MWAA) or AWS Step Functions. Cross-Cutting ● Lead code reviews, establish engineering best practices, and mentor junior/mid-level data engineers. ● Integrate data pipelines with CI/CD tools (Git, GitHub Actions, Jenkins, CodePipeline) for automated deployment. ● Monitor and optimize cloud costs across Databricks, Snowflake, and AWS workloads. ● Translate business requirements into technical data platform solutions with clear documentation. ◆ Required Skills & Qualifications Core Engineering ● 10+ years of Data Engineering experience; expert-level hands-on across Databricks, Snowflake, and AWS. ● Proficiency in Python and SQL; strong command of PySpark and/or Scala for distributed data processing. ● Deep understanding of distributed computing, data lake architecture, and cloud-native design patterns. Databricks ● 4+ years with Databricks: Delta Lake, Delta Live Tables, Databricks Workflows, Unity Catalog, and Spark optimization. ● Experience with Databricks on AWS, Azure, or GCP — cluster tuning, auto-scaling, and cost controls. Snowflake ● 4+ years with Snowflake: virtual warehouses, micro-partitioning, Snowpark, Streams, Tasks, and Snowpipe. ● Hands-on with Snowflake data modeling (dimensional, Data Vault), RBAC, and compliance frameworks. ● Experience with ETL/ELT tools: dbt, Fivetran, Matillion, or Informatica integrated with Snowflake. Data Quality ● 3+ years building DQ frameworks using Great Expectations, Soda Core, Apache Griffin, or Deequ. ● Experience with observability platforms: Monte Carlo, Acceldata, or Atlan; knowledge of OpenLineage/Marquez. ● Understanding of data governance standards, GDPR/CCPA compliance, and regulatory requirements. AWS ● 5+ years on AWS data services: S3, Glue, EMR, Redshift, Athena, Kinesis, MSK, and Lake Formation. ● Hands-on with Terraform or AWS CDK; familiarity with AWS IAM, VPC, and security best practices. ● Experience with MWAA (Airflow), Step Functions, or equivalent orchestration tools. ◆ Good to Have ● Databricks Certified Data Engineer Professional or Associate. ● SnowPro Core or SnowPro Advanced Architect certification. ● AWS Certified Data Engineer – Associate or Solutions Architect – Professional. ● Experience with Apache Iceberg or Hudi on S3; knowledge of real-time lakehouse patterns. ● Exposure to ML pipeline orchestration: MLflow, SageMaker, or Vertex AI. ● Familiarity with data mesh architectures, federated governance, and data contracts.