Data modelling
Why it matters
Data modelling decides whether a question like "what is revenue this month" gets one answer or three. When every dashboard reads raw data directly, each one rebuilds the logic slightly differently. A model settles it once. Without one, a business spends hours arguing about whose number is right instead of acting on any of them.
It matters more with AI and automation. An agent asked to report on revenue will use whichever definition it finds first, so the definitions have to be findable and unambiguous.
How to apply it
- Model a handful of core entities the business actually runs on, not every table that exists.
- Decide the grain of each table: one row per what? One row per customer, per subscription or per invoice line.
- Write business rules once inside the model and give each a plain-language name.
- Test the model with simple checks, such as no duplicate customer IDs and no invoice without a customer, and keep definitions under version control so changes are reviewable.
- Point every dashboard and report at the model, never at the raw source tables.
- Revisit definitions when the business changes, such as a new plan or region.
What it is
Raw exports and event logs are full of duplicates, odd field names and inconsistent values. A data model sits on top and describes the business in a few core tables, such as customers, subscriptions and invoices. It also states how they connect: one customer can have several subscriptions, and each subscription produces many invoices.
A model also holds the rules that give numbers their meaning. What counts as an active customer? When does a trial become a customer? How is a refund treated? These rules are written once, inside the model, rather than rebuilt in every report.
Common mistakes
- Modelling every table that exists instead of the few the business runs on.
- Leaving the grain of a table unclear, so a join silently multiplies rows and inflates revenue.
- Writing rules inside individual dashboards, which recreates the many versions of one number.
- Skipping tests, so a change in the source system breaks reports unnoticed.
- Naming columns for the tool they came from, not for what they mean.
- Treating it as a one-off project. A model needs an owner and a review when plans or regions change.