Watching an ingestion job's health and progress
Monitoring data ingestion in Microsoft Fabric involves tracking pipeline runs, dataflow refreshes, and streaming ingestion through built-in monitoring tools to ensure data arrives reliably and on schedule. Fabric provides centralized visibility into ingestion activities via the Monitoring Hub and item-specific run history, enabling engineers to detect failures, latency, and performance bottlenecks early.
Must-know
- The Monitoring Hub in Fabric provides a unified view across all workspaces showing status, start/end times, and duration for pipeline runs, dataflow refreshes, and other item activities the user has access to.
- Each Data Factory pipeline in Fabric has a Run history view showing individual activity-level status, inputs/outputs, and error messages, which is essential for diagnosing failed ingestion runs.
- Fabric pipelines support setting up alerts and notifications (via Activator or Data Activator integration) that trigger on specific conditions, such as pipeline failure or duration thresholds, allowing proactive monitoring rather than manual checking.
- For Eventstream and real-time ingestion, Fabric provides a visual monitoring canvas showing data flow, throughput, and connection health between sources, transformations, and destinations.
- Pipeline run and activity metadata can be queried programmatically via REST APIs, enabling custom monitoring dashboards or integration with external alerting systems beyond the native UI.
- A common gotcha is that Monitoring Hub only shows items the signed-in user has permission to view, so workspace admins may need broader role assignments to get full visibility into ingestion across teams.
- Dataflow Gen2 refresh history shows detailed step-by-step execution and duration per query/table, which helps pinpoint which specific data source or transformation is slowing down or failing during ingestion.
A data engineer at a retail company uses a Data Factory pipeline in Microsoft Fabric to run a Copy activity that ingests sales transactions into a lakehouse every night. The engineer has access to several workspaces and wants a single, consolidated view that shows the recent status, start time, and duration of pipeline runs across all of those workspaces, without opening each pipeline individually. Which feature should the engineer use?
What you have tried across DP-700's objectives, not a readiness score.
Implement and manage an analytics solution
- Tuning a workspace's Spark compute defaults and pool sizing
- Grouping and governing workspaces with a Fabric domain
- Setting per-workspace defaults for OneLake storage
- Standing up an Airflow job runtime inside a workspace
- Connecting a workspace to a Git repository
- Managing schema changes with a database project
- Promoting Fabric items across environments with a deployment pipeline
- Granting and restricting access at the workspace level
- Locking down who can open a single Fabric item
- Layering row, column, object, and file-level security rules
- Hiding sensitive column values behind a dynamic mask
- Classifying Fabric items with a sensitivity label
- Marking a trusted item as promoted or certified
- Reading a Fabric audit log to see who did what
- Securing data at the OneLake storage layer
- Picking the right build tool among a dataflow, a pipeline, and a notebook
- Kicking off a job on a schedule or in response to an event
- Chaining notebooks and pipelines together with parameters and dynamic expressions
Ingest and transform data
- Deciding between a full reload and an incremental load
- Shaping source data ahead of a dimensional-model load
- Landing a continuous stream of data into storage
- Matching a workload to the right Fabric data store
- Picking a transformation tool from dataflows, notebooks, KQL, or T-SQL
- Linking to external data without copying it via a OneLake shortcut
- Keeping a source database continuously replicated into Fabric
- Moving data into Fabric with a data pipeline
- Writing transform logic in PySpark, SQL, or KQL
- Flattening related tables into one wide, denormalized shape
- Rolling records up with group-by aggregations
- Dealing with duplicate rows, gaps, and data that arrives late
- Selecting the right engine for a real-time workload
- Weighing storage-in-place against a linked shortcut for a Real-Time Intelligence table
- Weighing an accelerated shortcut against a standard one for query speed
- Routing and reshaping live events with an Eventstream
- Handling a continuous flow of records with Spark's structured streaming
- Querying and reshaping event data with KQL
- Aggregating a stream over sliding or tumbling time windows
Monitor and optimize an analytics solution
- Watching an ingestion job's health and progress
- Watching a transformation job's health and progress
- Tracking whether a semantic model's refresh actually succeeded
- Setting up an alert to catch a failure early
- Tracking down why a pipeline run failed and fixing it
- Diagnosing why a dataflow run failed
- Debugging a notebook run that failed
- Troubleshooting a misbehaving Eventhouse
- Troubleshooting a misbehaving Eventstream
- Debugging a T-SQL statement that failed
- Fixing a broken or unreachable shortcut
- Speeding up a Lakehouse table with maintenance operations
- Making a slow pipeline run faster
- Tuning a Fabric warehouse for faster queries
- Improving throughput on real-time streaming components
- Tuning a Spark job to run faster and cheaper
- Making a slow query run faster
Coverage checked against the published exam guide on Aug 11, 2026.
These are independent practice questions, written against this certification's published exam guide. They are not the certification vendor's own questions, and not the real exam.