Data Engineering

Game Data Pipeline

RedshiftSQLdbtPrefectS3Python

An ingestion pipeline that pulls games metadata from the public IGDB API and shapes it into clean tables that feed a recommendation engine. I refactored it from row-by-row Pandas processing to set-based SQL running inside the warehouse, and moved the derived tables to dbt models — improving ingestion stability and data freshness.

Games-data ingestion and serving pipeline Ingest path: the public IGDB API is pulled by a scheduled flow into raw tables in Redshift. Serve path: SQL and dbt models curate the raw tables, payloads are exported to S3, and a recommendation engine consumes them. INGEST SERVE IGDB API public games data Ingestion flow scheduled with Prefect Redshift — raw one table per endpoint Transform SQL + dbt models Payloads exported to S3 Recommendation engine consumes curate
Raw endpoints land once; curated tables and payloads are derived in-warehouse and served to the recommendation engine.

The refactor

Outcome

The curated tables become the payloads a recommendation engine reads. Pushing the heavy work into SQL and standardizing on dbt made ingestion more stable and kept the data the recommender sees fresher.

Back to projects