Data Engineering
Game Data Pipeline
RedshiftSQLdbtPrefectS3Python
An ingestion pipeline that pulls games metadata from the public IGDB API and shapes it into clean tables that feed a recommendation engine. I refactored it from row-by-row Pandas processing to set-based SQL running inside the warehouse, and moved the derived tables to dbt models — improving ingestion stability and data freshness.
The refactor
- Pandas → SQL — replaced in-memory dataframe transforms with set-based SQL executed in Redshift, so transformation scales with the warehouse instead of a single worker's memory.
- Raw-first ingestion — each API endpoint lands in its own raw table; curation happens downstream, so a transform change never requires re-pulling the source.
- Materialized views → dbt — moved derived tables into dbt models for versioned, testable transformations and more predictable freshness.
- Scheduled flows — ingestion runs as scheduled Prefect deployments.
Outcome
The curated tables become the payloads a recommendation engine reads. Pushing the heavy work into SQL and standardizing on dbt made ingestion more stable and kept the data the recommender sees fresher.