Data Engineering
Self-Hosted Data Lakehouse
A self-hostable data platform — notebooks, a governed catalog, SQL warehouses, jobs, elastic Spark compute, and a natural-language AI/BI assistant — that runs on any Kubernetes cluster you control. It runs entirely on-premise (or in your own cloud VPC) with no dependency on a SaaS control plane.
Approach
Where strong open source already exists, the platform integrates it rather than reinventing it: Apache Polaris for the governed catalog, Delta Lake and Iceberg for table formats, Spark on the Kubeflow Spark Operator for compute, Trino for interactive SQL, Open Policy Agent for authorization, MinIO for object storage, and PostgreSQL for metadata.
What I built
- Control Plane — a FastAPI service that manages workspaces, compute (Spark clusters and SQL warehouses), jobs, notebooks, and a catalog proxy, driving the Kubernetes API and the governed catalog underneath.
- Natural-language assistant — a conversational AI/BI assistant that turns natural language into SQL against the governed catalog, with a pluggable model layer (Anthropic, OpenAI, Azure OpenAI, or local Ollama/vLLM).
- Workspace UI — a Next.js single-pane workspace with a catalog explorer, SQL editor, notebooks, compute, jobs, and the assistant chat.
How it was built
I built the platform end to end with an agentic-coding workflow — AI coding agents driving architecture, implementation, and iteration under close review — spanning a FastAPI control plane, a Spark and Trino compute layer on Kubernetes, and a Next.js workspace, kept coherent through tight specifications and continuous verification.