Keeping a source database continuously replicated into Fabric
Mirroring in Microsoft Fabric provides continuous, low-latency replication of data from operational databases and analytical stores directly into OneLake, without building or managing ETL pipelines. It creates a read-only replica in Delta Parquet format that stays automatically in sync with the source, enabling near real-time analytics across Fabric engines.
Must-know
- Fabric supports mirroring for sources such as Azure SQL Database, Azure Cosmos DB, Azure SQL Managed Instance, Snowflake, and on-premises SQL Server (via on-premises data gateway), plus Fabric database mirroring for Fabric-hosted databases.
- Mirrored data lands automatically in OneLake as Delta tables, so it can be queried directly by SQL analytics endpoints, Spark notebooks, Power BI, and other Fabric engines without duplication or additional copy steps.
- Mirroring is intended for full database or table-level replication with minimal configuration (typically enabling mirroring and selecting objects), not for row-level filtering, custom transformations, or incremental logic: use Dataflows or pipelines when transformation is required.
- Replication is near real-time (change data capture-based) and the mirrored replica is read-only; writes must occur at the source, and the source database must meet specific prerequisites (e.g., supported SKU tier, CDC enabled where applicable).
- Mirroring has no dedicated compute or Fabric capacity billing for the replication/storage of mirrored data itself in current Microsoft pricing guidance, making it a cost-effective way to bring source systems into OneLake, but this benefit only applies to items explicitly documented as eligible mirrored sources.
- Monitoring mirroring status (initial sync, replicating, or errors) is done in the Fabric portal via the mirrored database item's monitoring page, and common gotchas include network/firewall access to the source and unsupported data types or schema changes at the source disrupting replication.
A data engineer needs to bring change data from an on-premises PostgreSQL database into Fabric using Open Mirroring, since PostgreSQL is not one of the natively supported mirroring sources. The engineer will write a custom application to publish the changes. According to the Open Mirroring landing zone specification, what must this application do?
What you have tried across DP-700's objectives, not a readiness score.
Implement and manage an analytics solution
- Tuning a workspace's Spark compute defaults and pool sizing
- Grouping and governing workspaces with a Fabric domain
- Setting per-workspace defaults for OneLake storage
- Standing up an Airflow job runtime inside a workspace
- Connecting a workspace to a Git repository
- Managing schema changes with a database project
- Promoting Fabric items across environments with a deployment pipeline
- Granting and restricting access at the workspace level
- Locking down who can open a single Fabric item
- Layering row, column, object, and file-level security rules
- Hiding sensitive column values behind a dynamic mask
- Classifying Fabric items with a sensitivity label
- Marking a trusted item as promoted or certified
- Reading a Fabric audit log to see who did what
- Securing data at the OneLake storage layer
- Picking the right build tool among a dataflow, a pipeline, and a notebook
- Kicking off a job on a schedule or in response to an event
- Chaining notebooks and pipelines together with parameters and dynamic expressions
Ingest and transform data
- Deciding between a full reload and an incremental load
- Shaping source data ahead of a dimensional-model load
- Landing a continuous stream of data into storage
- Matching a workload to the right Fabric data store
- Picking a transformation tool from dataflows, notebooks, KQL, or T-SQL
- Linking to external data without copying it via a OneLake shortcut
- Keeping a source database continuously replicated into Fabric
- Moving data into Fabric with a data pipeline
- Writing transform logic in PySpark, SQL, or KQL
- Flattening related tables into one wide, denormalized shape
- Rolling records up with group-by aggregations
- Dealing with duplicate rows, gaps, and data that arrives late
- Selecting the right engine for a real-time workload
- Weighing storage-in-place against a linked shortcut for a Real-Time Intelligence table
- Weighing an accelerated shortcut against a standard one for query speed
- Routing and reshaping live events with an Eventstream
- Handling a continuous flow of records with Spark's structured streaming
- Querying and reshaping event data with KQL
- Aggregating a stream over sliding or tumbling time windows
Monitor and optimize an analytics solution
- Watching an ingestion job's health and progress
- Watching a transformation job's health and progress
- Tracking whether a semantic model's refresh actually succeeded
- Setting up an alert to catch a failure early
- Tracking down why a pipeline run failed and fixing it
- Diagnosing why a dataflow run failed
- Debugging a notebook run that failed
- Troubleshooting a misbehaving Eventhouse
- Troubleshooting a misbehaving Eventstream
- Debugging a T-SQL statement that failed
- Fixing a broken or unreachable shortcut
- Speeding up a Lakehouse table with maintenance operations
- Making a slow pipeline run faster
- Tuning a Fabric warehouse for faster queries
- Improving throughput on real-time streaming components
- Tuning a Spark job to run faster and cheaper
- Making a slow query run faster
Coverage checked against the published exam guide on Aug 11, 2026.
These are independent practice questions, written against this certification's published exam guide. They are not the certification vendor's own questions, and not the real exam.