Skip to content
Data engineering

From scattered sources to one warehouse that adds up

I bring sources together into a modelled, tested data warehouse: pipelines that run every day, figures that are the same for every department, and reporting nobody has to assemble by hand anymore.

That's the work I did as a data engineer at the Netherlands Aerospace Centre and as a data analyst at Politie Noord-Holland — two environments where a wrong figure has real consequences. The projects on this site show the same approach, with the code in the open.

Approach

What a team can expect from me

01

Model before building

Grain and dimensions on paper first, SQL second. A documented star schema stays understandable after I've moved on.

02

Tests on every change

Not-null, uniqueness and relationships as tests in the pipeline, with CI running them on every pull request. A mistake surfaces before it reaches a dashboard.

03

Batch and streaming

Daily ELT where that's enough, Kafka and windows where figures have to be right within minutes — including a dead-letter queue and replay for the failure path.

04

Explainable to the business

A warehouse isn't finished until the people using it understand it. Documentation and handover are part of the work, not an afterthought.

Experience

Where I did this

2024

Data Analyst

Politie Noord-Holland

Data analysis and automation within criminal law, under hard privacy and information-security requirements. Pipelines where privacy by design is the starting point, not an appendix.

2023

Data Engineer

Netherlands Aerospace Centre (NLR)

A wide range of data projects in a research environment where the figures get checked — and where the pipelines behind them have to hold up to that.

Since 2020

Instructor

Udemy

Courses on data and technology for people without a technical background: 241 students with a 4.2 rating. The same skill a team needs at every handover.

Works with
Data modeling
Kimball / star schema
dbt
SQL
Python
ELT pipelines
Kafka
Streaming
PostgreSQL
DuckDB
Docker
CI/CD
Azure
Power BI
Principles
  • One definition per figure, the same for every department.

  • Mistakes surface in the pipeline, not in the dashboard.

  • Everything reproducible: from zero with one command.

Projects

Projects that show it

Two own projects, from raw source to working result — with the full approach and the code in the open.

Own project

Open data warehouse: a Kimball star schema on RDW and CBS

End-to-end ELT platform on Dutch open data: Python ingestion, dbt transformations and a documented star schema — with tests and CI on every pull request.

dbt
SQL
Python
DuckDB
Data Engineering
Own project · 3 weeks

Real-time public transport streaming pipeline on Kafka

Every vehicle in Dutch public transport, streamed live — from GTFS-realtime feed to a map of the whole country, with the correctness guarantees to go with it: per-partition watermarks, at-least-once with idempotent upserts, a dead-letter queue and replay from the raw archive.

Kafka
Streaming
Python
Postgres
Docker

A warehouse that adds up, for your team?

I'm looking for a team where I can do this work properly. A conversation is easily arranged.