Diagnosing why a dataflow run failed
Dataflow Gen2 errors surface at multiple stages (authoring/Power Query evaluation, refresh execution, and data destination writes), and troubleshooting requires checking the refresh history, query diagnostics, and step-level error details in the Power Query editor. Understanding whether an error is a query folding issue, staging/compute issue, credential/gateway issue, or destination schema mismatch is key to resolving it efficiently. Fabric's monitoring hub and dataflow refresh history are the primary tools for diagnosing failures.
Must-know
- Step-level errors in the Power Query editor show a red exclamation icon and error message per cell/step, letting you isolate exactly which transformation caused the failure rather than debugging the whole query.
- The Refresh History (accessible from the dataflow item or Monitoring Hub) shows per-refresh status, duration, and detailed error messages, including whether the failure occurred during query evaluation, staging, or data destination write.
- Data destination errors (e.g., writing to a Lakehouse, Warehouse, or Azure SQL Database) often stem from schema drift, incompatible data types, or missing target permissions, and require checking the 'data destination settings' mapping for the affected table.
- Errors related to credentials or gateways (e.g., 'Unable to connect' or 'Credentials are invalid') require verifying the connection's authentication method and, for on-premises sources, confirming the On-premises Data Gateway is online and correctly configured.
- Query folding failures don't throw explicit errors but cause performance degradation; use the Query Diagnostics feature to identify steps that break folding and slow down refreshes.
- Staging-related errors can occur when the internal staging Lakehouse (used by Dataflow Gen2 with CI/CD or default destinations) hits capacity or throttling limits, which may require checking workspace capacity metrics in the Fabric Capacity Metrics app.
A Dataflow Gen2 loads transformed data into a Fabric Lakehouse table configured as the query's data destination. After the source system adds a new column, the next scheduled refresh fails with an error stating that the destination schema does not match the query output. The data engineer needs to resolve the current failure by aligning the destination table with the query's columns. What should the engineer do?
What you have tried across DP-700's objectives, not a readiness score.
Implement and manage an analytics solution
- Tuning a workspace's Spark compute defaults and pool sizing
- Grouping and governing workspaces with a Fabric domain
- Setting per-workspace defaults for OneLake storage
- Standing up an Airflow job runtime inside a workspace
- Connecting a workspace to a Git repository
- Managing schema changes with a database project
- Promoting Fabric items across environments with a deployment pipeline
- Granting and restricting access at the workspace level
- Locking down who can open a single Fabric item
- Layering row, column, object, and file-level security rules
- Hiding sensitive column values behind a dynamic mask
- Classifying Fabric items with a sensitivity label
- Marking a trusted item as promoted or certified
- Reading a Fabric audit log to see who did what
- Securing data at the OneLake storage layer
- Picking the right build tool among a dataflow, a pipeline, and a notebook
- Kicking off a job on a schedule or in response to an event
- Chaining notebooks and pipelines together with parameters and dynamic expressions
Ingest and transform data
- Deciding between a full reload and an incremental load
- Shaping source data ahead of a dimensional-model load
- Landing a continuous stream of data into storage
- Matching a workload to the right Fabric data store
- Picking a transformation tool from dataflows, notebooks, KQL, or T-SQL
- Linking to external data without copying it via a OneLake shortcut
- Keeping a source database continuously replicated into Fabric
- Moving data into Fabric with a data pipeline
- Writing transform logic in PySpark, SQL, or KQL
- Flattening related tables into one wide, denormalized shape
- Rolling records up with group-by aggregations
- Dealing with duplicate rows, gaps, and data that arrives late
- Selecting the right engine for a real-time workload
- Weighing storage-in-place against a linked shortcut for a Real-Time Intelligence table
- Weighing an accelerated shortcut against a standard one for query speed
- Routing and reshaping live events with an Eventstream
- Handling a continuous flow of records with Spark's structured streaming
- Querying and reshaping event data with KQL
- Aggregating a stream over sliding or tumbling time windows
Monitor and optimize an analytics solution
- Watching an ingestion job's health and progress
- Watching a transformation job's health and progress
- Tracking whether a semantic model's refresh actually succeeded
- Setting up an alert to catch a failure early
- Tracking down why a pipeline run failed and fixing it
- Diagnosing why a dataflow run failed
- Debugging a notebook run that failed
- Troubleshooting a misbehaving Eventhouse
- Troubleshooting a misbehaving Eventstream
- Debugging a T-SQL statement that failed
- Fixing a broken or unreachable shortcut
- Speeding up a Lakehouse table with maintenance operations
- Making a slow pipeline run faster
- Tuning a Fabric warehouse for faster queries
- Improving throughput on real-time streaming components
- Tuning a Spark job to run faster and cheaper
- Making a slow query run faster
Coverage checked against the published exam guide on Aug 11, 2026.
These are independent practice questions, written against this certification's published exam guide. They are not the certification vendor's own questions, and not the real exam.