How Did You Build That Pipeline?
It’s the question that comes up in every data engineering interview, every architecture review, every coffee chat with someone curious about the work. How did you build that pipeline? And the answer most people reach for is a list of tools: “Spark for processing, Airflow for orchestration, Parquet on S3, the usual.”
That answer tells you almost nothing. Tools are interchangeable and they change every two years. What doesn’t change is the order of thinking that turns a pile of raw data into something a business can trust. The engineers I’ve learned the most from can sketch that order on a napkin without naming a single product. That sketch - the overview - is the actual skill. This post is my attempt to write it down.
The reframe that changes everything: design backward, build forward
Here’s the mistake that quietly defines junior work: starting from the data. “I have this table of sensor readings - what can I do with it?” It feels productive. It’s the wrong direction.
Data flows forward, from source to output. But you design backward, from output to source. You start at the destination - what decision does this pipeline serve? - and reason your way back to the source, picking up requirements as you go. The destination shapes every choice upstream of it: how granular the data needs to be, how fresh, how correct. Get the destination wrong and you’ll build a beautiful pipeline that answers a question nobody asked.
So when I describe the stages below in forward order (because that’s how the data moves), keep the backward arrow in mind: every stage exists to serve the one after it, and ultimately to serve the consumer at the end.
The map
Strip away the tools and almost every batch pipeline is the same shape:
source -> raw -> clean -> model -> serve, with orchestration wrapped around it to run repeatably, and data quality threaded through it to know when it’s wrong.
Five nouns for the flow, two for what makes it production-grade. Let me walk them.
1. Define the output and its requirements
This is the most important stage and the one most often skipped. Before anything else: what question or decision does this pipeline serve? From that, three things fall out that constrain the entire design - the grain (per-second readings, or daily rollups?), the freshness (does this need to be real-time, or is a nightly batch fine?), and the level of correctness the use case actually demands.
These aren’t bureaucratic checkboxes. A dashboard that informs a quarterly review and a system that triggers automated alerts have completely different freshness and correctness budgets, and that difference ripples through every stage downstream. Skip this and you’re guessing.
2. Understand the source
This is the stage that separates levels, and it happens before you write a line of transformation code. What is the source, really? How does it deliver - batch files, change-data-capture, an event stream? What are its semantics - does it guarantee delivery at-least-once, meaning you should expect duplicates? What’s the schema, the volume, the typical lateness? Are the entities immutable (an event that happened and is fixed forever) or mutable (an order whose status keeps changing)?
Every one of those properties decides a design choice further down. At-least-once delivery means you’ll need deduplication. Mutable entities by key push you toward merge instead of overwrite. A source that’s routinely two days late dictates how far back your reprocessing has to reach. The junior move is to skip this and discover these properties the hard way, in production, as wrong numbers. The senior move is to interrogate the source until it has no surprises left.
3. Ingestion
Now you actually pull the raw data in - and the cardinal rule is to land it untouched. The raw layer is a faithful copy of the source, no transformations, no cleverness. It feels wasteful to store data you’re about to reshape anyway, but it’s the cheapest insurance you’ll ever buy: when your transformation logic turns out to be wrong (it will), the raw layer is what lets you rebuild everything downstream from scratch. You can’t reconstruct what you overwrote on the way in.
4. Storage and layout
How you physically organize the data. The Medallion pattern captures it well: bronze (raw) -> silver (cleaned, conformed) -> gold (modeled for consumption), each layer a deliberate step toward something a consumer can use. This is also where two decisions with outsized impact live: the file format (columnar formats like Parquet, because analytics reads columns, not rows) and partitioning.
Partitioning deserves a flag because it’s secretly a reliability decision, not just a performance one. Your partition is the unit of reprocessing - the chunk you can rebuild and overwrite in isolation without touching the rest. Partition by day, and reprocessing a bad day means rewriting one folder instead of the entire history. The layout you choose quietly determines how cheap it is to fix mistakes.
5. Transformation
The “T” everyone pictures when they think of data engineering, and - notice - it’s stage five, not stage one. Cleaning, deduplicating by natural key, conforming schemas, applying business logic, aggregating. This is where the reliability principles get enforced: where you make writes idempotent so reruns don’t multiply data, and where you collapse the duplicates that at-least-once delivery handed you. The transformations are only as trustworthy as the source understanding from stage 2 that informed them.
6. Orchestration and scheduling
This is what turns a set of scripts into a pipeline. How is it triggered - on a schedule, or by the arrival of data? How are dependencies between steps managed so they run in the right order? How does it process only what’s new (incremental) while still catching late-arriving data (a lookback window)? And critically: can it retry and backfill safely without corrupting anything?
The trigger choice has a hidden cost worth internalizing. A scheduler fires on the clock - it processes “yesterday” at 2am whether or not yesterday’s data has fully arrived. To compensate for that blindness, you reprocess a trailing window every run, just in case something showed up late. That window is the price you pay for triggering on time instead of on data. Event-driven pipelines, which fire because data arrived, largely avoid it. Either way, this stage is where individual steps become something that runs, unattended, every day, correctly.
7. Serving
Delivering the output to where it’s consumed - a table in the warehouse, an API, a dashboard. This closes the loop back to stage 1: the thing you serve has to match the grain, freshness, and shape the consumer needed. If stages 1 and 7 don’t agree, the work in between was aimed at the wrong target.
8. Data quality and observability
This isn’t the last step - it’s a layer woven through all the others, and it’s the second thing that separates levels. Data contracts, freshness SLAs, anomaly checks, lineage, alerting, cost monitoring. Its job is to catch the failures that don’t announce themselves.
Because that’s the thing about data pipelines: the dangerous bugs don’t crash. The job stays green, the schedule holds, the dashboard refreshes - and the numbers are quietly wrong. Without an observability layer actively checking, a pipeline that’s broken looks identical to one that’s working. This stage is the difference between finding out you’re wrong and a stakeholder finding out for you, three reports later.
What the map is really telling you
Look at where the two “separator” stages sit. Stage 2 (understand the source) and stage 8 (data quality) are the ones beginners skip - they jump straight into stages 3 through 5, the part that feels like real engineering, the writing of transformation code. They get a pipeline that runs. They don’t get a pipeline they can trust.
The progression from junior to senior isn’t about knowing more transformations. It’s about expanding outward from the middle: first learning to deeply understand the source you’re consuming, then learning to instrument the whole thing so it tells you when it lies. The code in the middle is necessary, but it’s the least transferable part - it’s the bookends that hold up.
So the next time someone asks how you built a pipeline, resist the tool list. The honest answer is a way of thinking: I started from what the output needed to be, I interrogated the source until I understood exactly what I was dealing with, I kept a raw copy I could always rebuild from, I organized storage so mistakes were cheap to fix, I made the transformations idempotent and correct, I orchestrated it to run and recover on its own, and I instrumented it so it would tell me the moment it went wrong.
The tools are just how you spell that out this year. The thinking is the part that lasts.
Đó là câu hỏi xuất hiện trong mọi buổi phỏng vấn data engineering, mọi buổi review kiến trúc, mọi cuộc trò chuyện bên tách cà phê với ai đó tò mò về công việc này. Bạn đã xây pipeline đó như thế nào? Và câu trả lời mà hầu hết mọi người với tới là một danh sách công cụ: “Spark để xử lý, Airflow để orchestration, Parquet trên S3, những thứ thông thường.”
Câu trả lời đó không nói cho bạn biết gần như điều gì. Các công cụ có thể thay thế cho nhau và chúng thay đổi mỗi hai năm. Thứ không thay đổi là trật tự tư duy biến một đống dữ liệu thô thành thứ mà một doanh nghiệp có thể tin tưởng. Những kỹ sư tôi học được nhiều nhất có thể phác thảo trật tự đó lên một tờ giấy mà không cần đặt tên một sản phẩm nào. Bản phác thảo đó - cái tổng quan - mới là kỹ năng thực sự. Bài này là nỗ lực của tôi để viết nó xuống.
Sự tái định nghĩa thay đổi tất cả: thiết kế ngược, xây xuôi
Đây là sai lầm âm thầm định nghĩa công việc của junior: bắt đầu từ dữ liệu. “Tôi có bảng sensor readings này - tôi có thể làm gì với nó?” Nó có vẻ hiệu quả. Nhưng đó là hướng sai.
Dữ liệu chảy xuôi, từ source đến output. Nhưng bạn thiết kế ngược, từ output về source. Bạn bắt đầu từ điểm đích - quyết định nào mà pipeline này phục vụ? - và lý luận ngược về source, thu thập các yêu cầu theo đường đi. Điểm đích định hình mọi lựa chọn ở thượng nguồn của nó: dữ liệu cần chi tiết đến mức nào, tươi mới đến mức nào, chính xác đến mức nào. Sai điểm đích và bạn sẽ xây một pipeline đẹp đẽ trả lời câu hỏi mà không ai hỏi.
Vì vậy khi tôi mô tả các giai đoạn bên dưới theo thứ tự xuôi (vì đó là cách dữ liệu di chuyển), hãy nhớ mũi tên ngược: mỗi giai đoạn tồn tại để phục vụ giai đoạn sau nó, và cuối cùng là để phục vụ người tiêu thụ ở cuối.
Bản đồ
Gạt bỏ các công cụ và hầu hết mọi batch pipeline đều có cùng một hình dạng:
source -> raw -> clean -> model -> serve, với orchestration bọc xung quanh để chạy lặp lại được, và data quality xuyên suốt để biết khi nào nó sai.
Năm danh từ cho dòng chảy, hai danh từ cho thứ làm nó đủ tiêu chuẩn production. Để tôi đi qua từng cái.
1. Xác định output và các yêu cầu của nó
Đây là giai đoạn quan trọng nhất và cũng là giai đoạn bị bỏ qua thường xuyên nhất. Trước tất cả mọi thứ: pipeline này phục vụ câu hỏi hay quyết định nào? Từ đó, ba thứ lộ ra và ràng buộc toàn bộ thiết kế - grain (số liệu theo từng giây, hay rollup theo ngày?), độ tươi mới (cái này cần real-time, hay batch nightly là đủ?), và mức độ chính xác mà use case thực sự đòi hỏi.
Đây không phải những checkbox hành chính. Một dashboard phục vụ review hàng quý và một hệ thống kích hoạt alert tự động có ngân sách về freshness và correctness hoàn toàn khác nhau, và sự khác biệt đó lan tỏa qua mọi giai đoạn downstream. Bỏ qua điều này và bạn chỉ đang đoán mò.
2. Hiểu source
Đây là giai đoạn phân tách cấp độ, và nó xảy ra trước khi bạn viết một dòng transformation code. Source thực sự là gì? Nó deliver thế nào - batch file, change-data-capture, một event stream? Ngữ nghĩa của nó là gì - nó có đảm bảo delivery at-least-once không, có nghĩa là bạn phải trông đợi duplicate? Schema là gì, volume bao nhiêu, lateness thông thường là bao lâu? Các entity có immutable không (một event đã xảy ra và cố định mãi mãi) hay mutable (một đơn hàng mà trạng thái vẫn tiếp tục thay đổi)?
Mỗi một trong những thuộc tính đó quyết định một lựa chọn thiết kế ở phía sau. At-least-once delivery có nghĩa là bạn sẽ cần deduplication. Mutable entity theo key đẩy bạn về phía merge thay vì overwrite. Một source thường xuyên trễ hai ngày quyết định bạn phải reprocess ngược về bao xa. Bước đi của junior là bỏ qua điều này và khám phá những thuộc tính này theo cách khó, trong production, dưới dạng những con số sai. Bước đi của senior là thẩm vấn source cho đến khi nó không còn bất ngờ nào nữa.
3. Ingestion
Giờ bạn thực sự kéo dữ liệu thô vào - và quy tắc cơ bản là land nó không chạm vào. Raw layer là bản sao trung thành của source, không transformation, không thủ thuật. Có vẻ lãng phí khi lưu trữ dữ liệu mà bạn sắp reshape lại, nhưng đây là bảo hiểm rẻ nhất bạn sẽ mua: khi transformation logic của bạn hóa ra sai (nó sẽ sai), raw layer là thứ cho phép bạn rebuild mọi thứ downstream từ đầu. Bạn không thể tái tạo thứ bạn đã overwrite trên đường vào.
4. Storage và layout
Cách bạn tổ chức dữ liệu một cách vật lý. Pattern Medallion nắm bắt điều này tốt: bronze (raw) -> silver (cleaned, conformed) -> gold (modeled để tiêu thụ), mỗi layer là một bước có chủ đích hướng đến thứ mà consumer có thể sử dụng. Đây cũng là nơi hai quyết định có tác động lớn nhất cư trú: file format (columnar format như Parquet, vì analytics đọc cột, không phải hàng) và partitioning.
Partitioning đáng được đánh dấu vì nó bí mật là một quyết định về reliability, không chỉ performance. Partition của bạn là đơn vị của reprocessing - khối bạn có thể rebuild và overwrite một cách độc lập mà không chạm vào phần còn lại. Partition theo ngày, và reprocessing một ngày xấu có nghĩa là ghi lại một folder thay vì toàn bộ lịch sử. Layout bạn chọn âm thầm xác định mức độ rẻ để khắc phục sai lầm.
5. Transformation
Chữ “T” mà mọi người hình dung khi họ nghĩ về data engineering, và - chú ý - đó là giai đoạn năm, không phải giai đoạn một. Cleaning, deduplicating theo natural key, conforming schema, áp dụng business logic, aggregating. Đây là nơi các nguyên tắc reliability được thực thi: nơi bạn làm cho việc ghi idempotent để rerun không nhân thêm dữ liệu, và nơi bạn collapse các duplicate mà at-least-once delivery đã trao cho bạn. Các transformation chỉ đáng tin cậy bằng sự hiểu biết về source từ giai đoạn 2 đã inform chúng.
6. Orchestration và scheduling
Đây là thứ biến một tập scripts thành một pipeline. Nó được trigger thế nào - theo lịch, hay bởi sự xuất hiện của dữ liệu? Các dependencies giữa các bước được quản lý thế nào để chúng chạy đúng thứ tự? Làm thế nào nó chỉ xử lý những gì mới (incremental) trong khi vẫn bắt được dữ liệu đến muộn (lookback window)? Và quan trọng hơn: nó có thể retry và backfill một cách an toàn mà không làm hỏng bất cứ thứ gì không?
Lựa chọn trigger có chi phí ẩn đáng được nội tâm hóa. Một scheduler fire theo đồng hồ - nó xử lý “hôm qua” lúc 2 giờ sáng dù dữ liệu của hôm qua có hoàn toàn đến hay không. Để bù đắp cho sự mù quáng đó, bạn reprocess một trailing window mỗi lần chạy, phòng khi có thứ gì đó đến muộn. Window đó là cái giá bạn trả cho việc trigger theo thời gian thay vì theo dữ liệu. Event-driven pipeline, fire vì dữ liệu đến, phần lớn tránh được điều đó. Dù thế nào, đây là giai đoạn nơi các bước riêng lẻ trở thành thứ chạy, không có người trông, mỗi ngày, đúng đắn.
7. Serving
Delivering output đến nơi nó được tiêu thụ - một bảng trong warehouse, một API, một dashboard. Điều này đóng vòng lặp trở lại giai đoạn 1: thứ bạn serve phải khớp với grain, freshness, và shape mà consumer cần. Nếu giai đoạn 1 và 7 không đồng ý, công việc ở giữa đã nhắm vào sai mục tiêu.
8. Data quality và observability
Đây không phải bước cuối cùng - đó là một layer được dệt xuyên suốt tất cả các bước khác, và đây là thứ thứ hai phân tách các cấp độ. Data contract, freshness SLA, anomaly check, lineage, alerting, cost monitoring. Công việc của nó là bắt những thất bại không tự thông báo.
Vì đó là điều về data pipeline: những bug nguy hiểm không crash. Job vẫn xanh, lịch vẫn giữ, dashboard refresh - và những con số âm thầm sai. Không có observability layer đang chủ động kiểm tra, một pipeline bị hỏng trông giống hệt một pipeline đang hoạt động. Giai đoạn này là sự khác biệt giữa việc biết ra bạn sai và một stakeholder biết ra cho bạn, ba báo cáo sau.
Bản đồ thực sự đang nói gì với bạn
Nhìn vào nơi hai giai đoạn “separator” ngồi. Giai đoạn 2 (hiểu source) và giai đoạn 8 (data quality) là những giai đoạn mà beginner bỏ qua - họ nhảy thẳng vào các giai đoạn 3 đến 5, phần trông như kỹ thuật thực sự, việc viết transformation code. Họ có được một pipeline chạy được. Họ không có được một pipeline có thể tin tưởng.
Sự tiến triển từ junior đến senior không phải là về việc biết nhiều transformation hơn. Đó là về việc mở rộng ra ngoài từ trung tâm: đầu tiên học cách hiểu sâu source bạn đang tiêu thụ, sau đó học cách instrument toàn bộ hệ thống để nó nói cho bạn biết khi nó nói dối. Code ở giữa là cần thiết, nhưng đó là phần ít có thể chuyển giao nhất - đó là những bookend giữ mọi thứ đứng vững.
Vì vậy lần sau khi ai đó hỏi bạn đã xây pipeline thế nào, hãy cưỡng lại danh sách công cụ. Câu trả lời trung thực là một cách tư duy: tôi bắt đầu từ những gì output cần phải là, tôi thẩm vấn source cho đến khi tôi hiểu chính xác mình đang xử lý gì, tôi giữ một bản sao raw mà tôi luôn có thể rebuild từ đó, tôi tổ chức storage để việc khắc phục sai lầm là rẻ, tôi làm cho các transformation idempotent và đúng đắn, tôi orchestrate nó để tự chạy và tự phục hồi, và tôi instrument nó để nó nói cho tôi biết ngay khi có gì đó sai.
Các công cụ chỉ là cách bạn đánh vần điều đó năm nay. Cách tư duy mới là phần tồn tại lâu dài.