Data pipeline

Definition
A data pipeline is the automated route data travels from where it is created to where it is used, cleaned and reshaped along the way.

Why it matters

Hand exports work once. Done every week, they cost time and add copy-paste errors, and two people end up reading two slightly different versions of the same number. A pipeline keeps every destination current without anyone remembering to click export.

The real risk is a pipeline that fails quietly. A broken sync is worse than no sync, because a team keeps trusting numbers that stopped updating days ago.

How to apply it

  • Start with the two or three sources that answer your most important question, not every tool you own.
  • Use a managed connector or an automation tool for common sources instead of writing a script for each one.
  • Schedule the run, and send an alert when a run fails or when no new rows arrive.
  • Keep the raw data exactly as it arrived, and do the cleaning as a separate, visible step.
  • Make a re-run safe: running the same job twice should never create duplicate rows.

What it is

A data pipeline is a set of steps that runs by itself. It collects data from a source, tidies it into a shape the destination can use, and delivers it. Three stages are common: extract (copy the data out of a source system), transform (fix formats, remove duplicates, join records) and load (write the result to its destination). The same stages appear under the names ETL / ELT, depending on whether the tidying happens before or after the data lands.

A webshop owner keeps orders in Shopify, ad spend in Google Ads and customer emails in a mailing tool. Each night a pipeline copies all three into one table. The next morning a single report shows spend, orders and revenue per channel, and nobody has exported a file.

Common mistakes

  • Building the pipeline before knowing which question the data must answer.
  • Overwriting raw data, so a mistake in the cleaning step can never be undone.
  • Having no owner, so nobody notices when a source changes its format.
Worked example

Suppose a webshop owner keeps orders on one platform, ad spend in an ad account and customer emails in a mailing tool. Every Monday someone exports three files, pastes them into a spreadsheet and fixes the date formats by hand. The export takes two hours, and twice a month two people produce different totals. The team builds a scenario in Make that runs each night. It reads yesterday's orders and ad spend, then writes both into one table with the same date format and one channel name for each source. A failed run is noticed at once, because the row count is checked each morning. In this example, the manual work drops from two hours a week to nothing, and the weekly report shows spend against orders for each channel before the first coffee of the day.

Tools in the example

Some links are affiliate links: we may earn a commission at no cost to you. It never decides a ranking. How we work with partners

  1. Article

    ETL / ELT

    The two patterns a pipeline follows when data is loaded into a warehouse.

  2. Article

    Data warehouse

    The usual destination a pipeline delivers into.

  3. Article

    Reverse ETL

    A pipeline running the other way, back out to everyday tools.

  4. Article

    Data Hygiene

    The habit of keeping the records a pipeline carries clean.

  5. Article

    Workflow automation

    The wider family of automated flows that a pipeline belongs to.