Ian Ronk
Ian Ronk
Data Lead · Data Engineer & System Architecture
Amsterdam, NL

Transforming
complex data
into insights.

I am Ian and I build and maintain production data systems and the analytics on top, from any complex data type: weekly high-volume web-scraped data, geospatial, graph or tabular data. I answer stakeholder questions using this data.

Focus areasData EngineeringSystem ArchitectureComplex Data ProductsAnalytics & ML

Hi, I'm Ian.
I create insights using data

I work as a Head of Data, where I build and maintain data systems, collecting, processing and structuring data for different data products, used by pension funds, bureaus of statistics and real estate investors with 10B AuM.

Over the course of my career I worked with many different types of data and tools, creating custom systems for the data problems at hand, talking to client stakeholders and creating innovative solutions.

Read the full bio
Ian Ronk
RoleData Lead · Data Engineer
BasedAmsterdam, NL
EducationMSc Bocconi · BSc AI, UvA
StackBash · Python · PostGIS/SQL · Airflow · Iceberg

Focus Areas

Working as Head of Data I wear many hats. My four focus areas are building the data infrastructure, setting up the systems itself, handling different types of data and running analytics on top.

§ 03.01

Data Engineering

Building and maintaining data pipelines and efficient storage, such as three years of weekly collection/scraping across 8 different sources for a real estate fund with 10B AuM and then structuring, cleaning and deduplicating this data automatically.

AirflowPythonETLMonitoringIceberg
§ 03.02

System Architecture

Setting up the platform underneath: server instances, PostGIS, distributed Iceberg compute and data processing servers, networking/VPN, APIs and security. Example: 13 different services run as production infrastructure to handle loads of complex data products.

Linux/BashResource Management & S3Networking & APIsDockerDistributed Compute
§ 03.03

Complex Data Products

Turning complex and unconventional data types into usable data for analytics: spatial and network data, document data and time series, owned end-to-end from raw data to insights.

SpatialGraphNoSQL/MongoDBOCRData Validation
§ 03.04

Analytics & ML

The analysis layer on top: nowcasting, ABM simulations, regressions, statistical methods and applied ML that ships: from hedonic price models, to image classification to statistically sound time series analyses.

RegressionTime SeriesSimulationsInsightsStatistical Analysis

Projects & papers.

Eight pieces of work across the four lanes: production systems, shipped products, and the research they make possible.

conversationlemmatizeSM-21d6d15dnext-day6,200 vocab · A1–C1
§ 04.01AI productPROJECT

LanguageBuddy: AI language tutor

A self-hosted AI language tutor for Dutch, Italian and Spanish, built on language learning research that includes a chat or voice-call LLM tutor and SM-2 repetition. Every mistake is captured into a spaced-repetition queue that drives the next day's exercises and uses a real-news scraper for reading learning. Follows the CEFR framework with 6,200+ vocabulary entries and helps users reach CEFR alignment faster.

FastAPILLMTTSSQLiteDocker
t=0t=10yattractiveness ↑affordability ↓parcels: AMS · UTRECHT · MILAN
§ 04.02MSc Thesis · BocconiRESEARCH

Gentrification agent-based model

My MSc Thesis research on gentrification, where I built an agent-based model of neighbourhood change: households interact with neighborhoods in the treated cities Amsterdam, Utrecht and Milan based on their own income, the affordability of the neighborhood and the attractiveness, based on elements such as streetview imagery 'beauty' classification, GTFS connectivity, greenery and neighborhood sentiment online. A two-tenure social-housing extension is in progress.

Python/GeoPandasLarge Vision ModelPostGISAgent-based Modeling
DAGingest726 GBguardtransformqueuecelery workers
§ 04.03EngineeringPROJECT

Research pipelines as production systems

Every personal research project I undertake uses Airflow 3 data pipelines that are hosted in this project. It can scale from a laptop to a multi-machine CeleryExecutor cluster. This project was built with idempotency guards, custom operators and agentic monitoring.

AirflowCeleryDockerCIDuckDB
p(sponsor)86% F10.54:326:05sentence-T5 → BiLSTM · 38,600 videos
§ 04.04NLP · sequence taggingPROJECT

SponsoredBye: sponsor-segment detection

A text-only sponsor-skipper for YouTube, built before YouTube Premium shipped one: sentence-T5 embeddings feed a BiLSTM sequence tagger that flags sponsored sentences and maps them back to timestamps, cutting segmentation error from 99% to 16% WindowDiff.

TensorFlowBiLSTMsentence-T5MongoDBHuggingface
SAM224pxResNet501/638.8 MB on-device
§ 04.05Mobile MLPROJECT

FishFinder: photo-to-species ID

A Flutter app that identifies 63 Dutch fish species from a photo, fully on-device, and fills a Pokédex-style FishDex as you catch them. The training pipeline: ~3,000 hand-annotated photos masked with Segment Anything Model, then a fine-tuned ResNet50 compressed to an 8.8 MB TFLite model for on-device use.

Flutter/DartTFLiteResNet50Segment Anything Model (SAM)Firebase
RF 97.5%33 featuresd(river)Δhimperv. 500m–5kmw.l.
§ 04.06BSc Thesis · UvARESEARCH

Predicting flooding risk from local features

My BSc thesis, looking at whether flood risk can be explained by local features instead of a black-box hydrodynamic simulation. Collected 33 features across ~45,000 European locations, such as ground imperviousness, ground type and distance to river, across 100GB+ of data and ran a Random Forest to find their relation. Findings: 97.5% on the binary 20-year flood question, with surrounding imperviousness and relative height doing most of the work.

scikit-learnRandom Forestraster dataGIS
945 min15 min38 markets · EU/NA/APACparcel score 0–100
§ 04.07Developed at KR&AKR&A

Connectivity & walkability scoring

A project led and conducted while working at KR&A: a saturation-validated connectivity and walkability score at parcel resolution, rolled out across 38 EU/NA/APAC markets, working with TBs of data and 100s of data sources.

PostGISS3Dijkstra & Network ScienceDistributed Compute
monthlyEurostat Qbase=10013 countries
§ 04.08Developed at KR&AKR&A

Monthly house-price index · 13 EU countries

A project conducted while working at KR&A: a web-scraping pipeline across 13 countries feeding a log-price hedonic regression; monthly indices tested as disaggregation indicators for Eurostat quarterly HPIs.

PythonScrapy/SeleniumPostGISMongoDBR

Papers.

Methods, data and results behind the projects. All papers →

Questions, ideas, roles? Get in touch.

ContactGitHub
Ian Ronk | Head of Data: data systems, analytics, and urban-dynamics research