Data pipeline
Why it matters
Hand exports work once. Done every week, they cost time and add copy-paste errors, and two people end up reading two slightly different versions of the same number. A pipeline keeps every destination current without anyone remembering to click export.
The real risk is a pipeline that fails quietly. A broken sync is worse than no sync, because a team keeps trusting numbers that stopped updating days ago.
How to apply it
- Start with the two or three sources that answer your most important question, not every tool you own.
- Use a managed connector or an automation tool for common sources instead of writing a script for each one.
- Schedule the run, and send an alert when a run fails or when no new rows arrive.
- Keep the raw data exactly as it arrived, and do the cleaning as a separate, visible step.
- Make a re-run safe: running the same job twice should never create duplicate rows.
What it is
A data pipeline is a set of steps that runs by itself. It collects data from a source, tidies it into a shape the destination can use, and delivers it. Three stages are common: extract (copy the data out of a source system), transform (fix formats, remove duplicates, join records) and load (write the result to its destination). The same stages appear under the names ETL / ELT, depending on whether the tidying happens before or after the data lands.
A webshop owner keeps orders in Shopify, ad spend in Google Ads and customer emails in a mailing tool. Each night a pipeline copies all three into one table. The next morning a single report shows spend, orders and revenue per channel, and nobody has exported a file.
Common mistakes
- Building the pipeline before knowing which question the data must answer.
- Overwriting raw data, so a mistake in the cleaning step can never be undone.
- Having no owner, so nobody notices when a source changes its format.