Skip to content
60 minute session

From Convenience to Control: Breaking Free from Big Tech Data Platforms

Almost every data team starts its day inside someone else’s data platform: identity, ingestion, storage, transformation, dashboards, and deployment. This session opens with an uncomfortable question, “who could still run their data pipelines on Monday if their hyperscaler account were blocked?”, and works backwards from there to explain how we quietly traded control for convenience, one layer at a time.

It then looks at what sovereignty actually means. Not “the data sits in an EU region”, but control: who can read the data, who can switch it off, whether you could rebuild it elsewhere and who makes the parts it runs on. Sovereignty turns out to be a dial rather than a switch, and different workloads, such as public content, internal operational data, privacy-sensitive data, safety-critical systems all belong at different settings.

The second half is practical. Starting from four design principles, the session builds up a reference architecture for a sovereign data platform out of open-source components such as Apache Iceberg on Parquet, an Iceberg REST catalogue, Trino and DuckDB, dlt and dbt-core, Airflow, Keycloak and self-hosted object storage, with a live demo of a working open-source lakehouse.

It closes honestly: owning your control plane also means owning the patching, the security operations and the on-call rotation. A sovereign platform is an HR project disguised as a technology challenge, and you are not buying a cheaper platform, you are buying the ability to get and stay in control.

Learning goals
  • Recognising how much of your daily data workflow depends on a single provider
  • Understanding how infrastructure abstraction, legislation such as the CLOUD Act and weak negotiating positions eroded control
  • Explaining what data sovereignty really means, and why data location alone is not sovereignty
  • Assessing the four layers of control (data, operations, technology and supply chain) for your own platform
  • Matching the level of sovereignty to the sensitivity and criticality of each workload
  • Applying the design principles behind a sovereign data platform, including why open-source is not automatically sovereign
  • Naming the open-source building blocks of a sovereign lakehouse reference architecture
  • Weighing the real costs of self-hosting, from security operations and patching to the people you need to run it
  • Identifying the first concrete steps towards regaining control of your data platform
Level
Intermediate
Topics
Data Sovereignty Open Source Data Platforms Data Privacy
Materials
Slides
All sessions