Onhires

Senior Data Engineer (AWS / Databricks)

Europe (remote) · 1 month ago
Tech Stack
AWSDatabricksSparkPySparkS3Delta LakePostgreSQLPythonSQL TerraformUnity CatalogDatabricks JobsDatabricks WorkflowsAWS DMSDebeziumdbtEKS
Language Requirements
English
Requirements
Senior Seniority
No Degree
Remote Policy

Remote

EU, Ukraine

Remote (EU/Ukraine) | Full-time

We’re hiring on behalf of our client — an international product company building and scaling a portfolio of subscription-based digital products for global markets.

The company is now building a centralized data platform that will bring together fragmented product, payments, marketing, and operational data across its portfolio. They are looking for a hands-on Senior Data Engineer to help establish Databricks on AWS, define the platform’s core engineering standards, and build reliable data products for analysts and business stakeholders.

This is a greenfield platform role with substantial technical ownership. You will not be joining a mature data environment with established patterns. You will help design those patterns, make foundational technical decisions, and create a repeatable approach for onboarding new products and data sources.

Why this role is interesting

  • Build a centralized data platform from an early stage rather than inherit a mature warehouse

  • Influence architecture, engineering standards, ingestion patterns, and governance

  • Solve a complex platform challenge involving distributed PostgreSQL databases across private AWS and EKS environments

  • Build the first portfolio-wide data models around payments, subscriptions, revenue, churn, LTV, and CAC

  • Work closely with Data, DevOps, Product, Backend Engineering, and business stakeholders

  • See a direct connection between your engineering work and key product and commercial decisions

What you’ll do

Build the Data Platform

  • Build and operate a Databricks-based data platform on AWS together with the Data and DevOps teams

  • Design and maintain Bronze, Silver, and Gold data layers using S3 and Delta Lake

  • Develop reusable ingestion patterns for PostgreSQL databases, S3, APIs, webhooks, and SaaS platforms

  • Build and manage production workflows using Databricks Jobs and Workflows

  • Contribute infrastructure changes through Terraform, Git, and pull-request-based workflows

  • Help establish platform standards, development patterns, and technical documentation

Build Reliable Data Pipelines

  • Implement incremental data loads, historical backfills, idempotent reprocessing, and schema-change handling

  • Design safe ingestion from multiple production PostgreSQL databases without creating unnecessary risk or load for source applications

  • Handle late-arriving updates, deletes, retries, and pipeline recovery

  • Build monitoring, freshness checks, reconciliation processes, and data-quality controls

  • Troubleshoot pipeline failures and data inconsistencies across multiple products and source systems

  • Optimize Databricks compute, SQL workloads, and storage for performance, reliability, and cost

Unify Product and Payments Data

  • Standardize fragmented product and payments data across the company’s portfolio

  • Build common analytical entities for users, subscriptions, transactions, renewals, refunds, and chargebacks

  • Normalize product-specific schemas into reliable source-of-truth models

  • Deliver trusted Gold datasets and data marts for Payments, Marketing, Product, Finance, and executive reporting

  • Support analytical use cases related to revenue, subscriptions, churn, LTV, CAC, product funnels, and attribution

Establish Governance and Engineering Standards

  • Contribute to Unity Catalog implementation and ongoing governance

  • Help manage groups, permissions, service principals, and data access patterns

  • Apply Git-based development, code review, CI/CD, testing, and documentation practices

  • Work with Product and Backend teams to understand source tables, relationships, and business logic

  • Help define repeatable patterns for onboarding new products and data sources

Target Platform

  • AWS

  • Databricks

  • Spark / PySpark

  • S3

  • Delta Lake

  • Unity Catalog

  • Databricks Jobs / Workflows

  • PostgreSQL

  • Python

  • SQL

  • Terraform

  • Git and CI/CD

What we’re looking for

  • Strong production experience in Data Engineering

  • Advanced Python and SQL skills

  • Hands-on production experience with Databricks and Spark/PySpark

  • Practical AWS experience, particularly with S3 and IAM

  • Experience ingesting data from PostgreSQL or other relational databases

  • Strong understanding of incremental pipelines, historical backfills, idempotency, retries, and reprocessing

  • Experience designing analytical data models and working with medallion architecture

  • Experience implementing data-quality checks, monitoring, reconciliation, and troubleshooting

  • Experience with Git-based development and CI/CD workflows

  • Ability to take ownership of complex data initiatives from design through production operation

  • Comfort working in a greenfield environment where standards, ingestion patterns, and models are still being defined

  • Ability to collaborate effectively with DevOps, Backend Engineering, Product, Analytics, and business stakeholders

Strong advantages

  • Experience with Terraform or another Infrastructure as Code tool

  • Hands-on experience with Unity Catalog

  • Experience with Databricks Jobs, Workflows, or Lakeflow

  • Understanding of AWS networking, VPCs, and EKS environments

  • Experience with CDC technologies such as AWS DMS or Debezium

  • Experience working with subscription and payments data

  • Familiarity with Stripe, Adyen, Solidgate, or other payment service providers

  • Experience integrating marketing or attribution data

  • Experience with dbt

  • Previous responsibility for defining data-platform standards or reusable engineering patterns

  • Experience building a data platform in a startup, scale-up, or other ambiguous environment

What success could look like

During your first stage in the role, you will help:

  • Establish the core AWS, Databricks, S3, Unity Catalog, and Terraform platform foundation

  • Productionize the first reusable end-to-end ingestion pattern

  • Onboard and unify payments data across multiple products

  • Deliver core Gold models for revenue, subscriptions, churn, and LTV

  • Create and document a repeatable approach for onboarding additional products

  • Put monitoring, reconciliation, CI/CD, and cost controls into production

What our client offers

  • Competitive compensation

  • Fully remote work with flexible working hours

  • 22 paid vacation days plus local public holidays

  • A modern engineering environment with contemporary technologies

  • The opportunity to shape a growing Data function and its technical foundations

  • Meaningful platform challenges with room to influence architecture and engineering practices

  • A collaborative, product-focused environment where data directly supports business decision