What is Reference DB?

The Reference DB is a structured data system designed to transform product data collected from different Canadian retailers into consistent, reliable, and reusable datasets.

Retailers do not provide product information in the same format. They may use different:

Reference DB provides a common processing framework that turns these different source datasets into standardized retailer-level reference datasets.

Reference DB is not simply a scraper and not simply a database. It is a pipeline that separates source data, source-specific adaptation, validation, identity/semantic normalization, classification, taxonomy, grouping, variants, nutrition, and scores/enrichment into distinct, testable stages.

Why do we need this system?

Suppose we collect products from several retailers — Walmart, Costco, Metro, Voilà. Each source arrives with different schemas, identifiers, names, categories, nutrition formats, and product structures.

For example, the same type of product might appear as:

Walmart:  Peanut Butter Smooth 500g
Costco:   Kirkland Signature Smooth Peanut Butter - 500 g
Voilà:    Compliments Smooth Peanut Butter 500 g

The source systems may also classify products differently — one retailer nests “Spreads” under Grocery → Pantry, another uses a flat “Pantry” category, and a third uses “Peanut Butter & Spreads” as its own category.

If downstream systems directly consume the original retailer data, every consumer has to understand every retailer's structure — duplicated logic, hard to maintain. Reference DB introduces a controlled semantic layer between source data and downstream consumers.

The Core Architecture