Data Engineering

Self-Hosted Data Lakehouse

KubernetesSparkTrinoIcebergDelta LakeApache PolarisFastAPINext.js

A self-hostable data platform — notebooks, a governed catalog, SQL warehouses, jobs, elastic Spark compute, and a natural-language AI/BI assistant — that runs on any Kubernetes cluster you control. It runs entirely on-premise (or in your own cloud VPC) with no dependency on a SaaS control plane.

Approach

Where strong open source already exists, the platform integrates it rather than reinventing it: Apache Polaris for the governed catalog, Delta Lake and Iceberg for table formats, Spark on the Kubeflow Spark Operator for compute, Trino for interactive SQL, Open Policy Agent for authorization, MinIO for object storage, and PostgreSQL for metadata.

What I built

How it was built

I built the platform end to end with an agentic-coding workflow — AI coding agents driving architecture, implementation, and iteration under close review — spanning a FastAPI control plane, a Spark and Trino compute layer on Kubernetes, and a Next.js workspace, kept coherent through tight specifications and continuous verification.

Back to projects