Dealing with duplicate rows, gaps, and data that arrives late
Real-world data pipelines in Microsoft Fabric must account for duplicates, nulls, and data that arrives after its expected processing window. Fabric provides tools across Dataflow Gen2, pipelines, and Spark notebooks to detect and remediate these quality issues before data lands in the lakehouse or warehouse. Choosing the right technique (dedup keys, watermarking, imputation) depends on the ingestion pattern (batch vs. streaming) and downstream analytical requirements.
1 · Learn the must-know
- Deduplication can be handled in Dataflow Gen2 using the 'Remove Duplicates' transformation, or in Spark/SQL using dropDuplicates(),
ROW_NUMBER() window functions, or MERGE/upsert logic on a business key to avoid inserting repeated rows. - Missing values can be handled via Dataflow Gen2's 'Fill Down/Fill Up' and 'Replace Values' transforms, or in Spark using fillna(), dropna(), or conditional imputation logic, depending on whether nulls should be defaulted, interpolated, or excluded.
- Late-arriving data in streaming scenarios (e.g., Eventstream or Spark Structured Streaming) is managed using watermarking, which defines how long the engine waits for delayed events before finalizing a windowed aggregation, trading completeness for latency.
- For batch pipelines, late-arriving dimension or fact records are typically handled with upsert/MERGE patterns keyed on business/natural keys plus a last-modified or ingestion timestamp, ensuring late updates overwrite or supplement existing lakehouse/warehouse rows rather than duplicating them.
- Delta Lake tables (the default table format in Fabric Lakehouse) natively support MERGE INTO, which is the standard mechanism for idempotently applying inserts/updates/deletes when reprocessing or handling late/duplicate data.
- A common gotcha: deduplication logic must define which record is 'correct' when duplicates exist (e.g., latest by timestamp) since simple distinct/removal without ordering can silently drop the wrong row.
2 · Check your understanding
An engineer builds a Fabric Dataflow Gen2 that ingests daily sales files from a data lake folder into a lakehouse table. Upstream systems occasionally reprocess a file, causing the same rows to appear more than once in the combined dataset. The dataflow must output a table with no duplicate rows. Which query step should the engineer add before loading the data?
What you have tried across DP-700's objectives, not a readiness score.
Implement and manage an analytics solution30-35% of the exam0 of 18 tried
Ingest and transform data30-35% of the exam0 of 19 tried
Monitor and optimize an analytics solution30-35% of the exam0 of 17 tried
3 · Keep going
Ready for more? Take a weighted mock or try free practice questions.