Skip to content

Picking the right build tool among a dataflow, a pipeline, and a notebook

Microsoft Fabric offers three primary tools for data ingestion and transformation: Dataflow Gen2 for low-code visual ETL, Pipelines for orchestration and copy activities at scale, and Notebooks for code-first transformations using Spark. Choosing the right tool depends on the persona (citizen integrator vs. data engineer), the complexity of transformation logic, data volume, and whether orchestration versus transformation is the primary goal.

Must-know

  • Dataflow Gen2 uses Power Query (M language) and is best suited for low-code, self-service data transformation by users comfortable with Power BI-style dataflows, but it can become slower and costlier than Spark for very large datasets.
  • Pipelines are primarily for orchestration (scheduling, chaining activities, and moving data at scale via Copy Activity) rather than complex row-level transformations; they can invoke Dataflows, Notebooks, or other pipelines as steps.
  • Notebooks (using PySpark, Scala, SQL, or R) provide the most flexibility and performance for complex, large-scale transformations and are the preferred choice for data engineers who need full control, version control (Git integration), and custom logic.
  • A common gotcha: Dataflow Gen2 output can be consumed by a Pipeline (e.g., as a source for further orchestration) or triggered from within a Pipeline, so the tools are often combined rather than mutually exclusive.
  • Notebooks scale better for very large or complex transformations because they leverage Spark's distributed compute, whereas Dataflow Gen2's Power Query engine can hit performance/cost limits at high volumes.
  • For simple scheduled copy/move operations with minimal transformation, Pipelines with Copy Activity are the most efficient and lowest-maintenance choice compared to spinning up a Notebook or Dataflow Gen2.
Check this objectiveFree · always available

A data engineer must build a transformation that calls a custom PySpark function to enrich streaming IoT sensor readings with results from an external REST API, and needs to develop the logic interactively, running individual cells and inspecting intermediate DataFrames before scheduling it for production. Which Fabric item should the engineer use to author this transformation?

Your objective map0 tried · 0 right · 54 untouched

What you have tried across DP-700's objectives, not a readiness score.

Coverage checked against the published exam guide on Aug 11, 2026.

These are independent practice questions, written against this certification's published exam guide. They are not the certification vendor's own questions, and not the real exam.