Linking to external data without copying it via a OneLake shortcut
OneLake shortcuts are references (not copies) that let you access data stored in another location (another OneLake path, ADLS Gen2, S3, or Dataverse) directly from a lakehouse or KQL database as if it were local. They enable unified, virtualized data access across workspaces and domains without duplicating storage or ETL, supporting a single copy of data (OneLake, one copy).
Must-know
- Shortcuts are created inside a lakehouse's Files or Tables area (or a KQL database) via the 'New shortcut' option, choosing an internal OneLake target or an external source like ADLS Gen2, Amazon S3, Google Cloud Storage, or Dataverse.
- A shortcut behaves like a symbolic link: reads go directly to the source data, so you always see the latest data without needing a copy or refresh job.
- Deleting a shortcut only removes the reference in OneLake; it does not delete the underlying source data, but deleting the source data can break the shortcut.
- Table shortcuts to folders containing Delta Lake-formatted data are automatically recognized as Delta tables and can be queried directly with SQL or Spark; shortcuts to non-Delta folders appear under Files.
- Cross-workspace shortcuts let you avoid data duplication and enable shared, governed access to a single source of truth, but the calling identity still needs permission on the underlying source (OneLake or external storage credentials/connection)
- Shortcuts support caching for performance on repeated reads from external sources, but permissions and authentication to the external system (e.g., SAS token, service principal, or workspace identity) must be correctly configured or the shortcut will fail to resolve.
A data engineer at a retail company needs the 'Sales' Delta table, which lives in the Tables section of the Finance team's lakehouse, to also appear inside her own lakehouse so her notebooks can query it directly. She wants any updates the Finance team makes to be visible immediately, and she does not want to duplicate the underlying Parquet files. What should she do?
What you have tried across DP-700's objectives, not a readiness score.
Implement and manage an analytics solution
- Tuning a workspace's Spark compute defaults and pool sizing
- Grouping and governing workspaces with a Fabric domain
- Setting per-workspace defaults for OneLake storage
- Standing up an Airflow job runtime inside a workspace
- Connecting a workspace to a Git repository
- Managing schema changes with a database project
- Promoting Fabric items across environments with a deployment pipeline
- Granting and restricting access at the workspace level
- Locking down who can open a single Fabric item
- Layering row, column, object, and file-level security rules
- Hiding sensitive column values behind a dynamic mask
- Classifying Fabric items with a sensitivity label
- Marking a trusted item as promoted or certified
- Reading a Fabric audit log to see who did what
- Securing data at the OneLake storage layer
- Picking the right build tool among a dataflow, a pipeline, and a notebook
- Kicking off a job on a schedule or in response to an event
- Chaining notebooks and pipelines together with parameters and dynamic expressions
Ingest and transform data
- Deciding between a full reload and an incremental load
- Shaping source data ahead of a dimensional-model load
- Landing a continuous stream of data into storage
- Matching a workload to the right Fabric data store
- Picking a transformation tool from dataflows, notebooks, KQL, or T-SQL
- Linking to external data without copying it via a OneLake shortcut
- Keeping a source database continuously replicated into Fabric
- Moving data into Fabric with a data pipeline
- Writing transform logic in PySpark, SQL, or KQL
- Flattening related tables into one wide, denormalized shape
- Rolling records up with group-by aggregations
- Dealing with duplicate rows, gaps, and data that arrives late
- Selecting the right engine for a real-time workload
- Weighing storage-in-place against a linked shortcut for a Real-Time Intelligence table
- Weighing an accelerated shortcut against a standard one for query speed
- Routing and reshaping live events with an Eventstream
- Handling a continuous flow of records with Spark's structured streaming
- Querying and reshaping event data with KQL
- Aggregating a stream over sliding or tumbling time windows
Monitor and optimize an analytics solution
- Watching an ingestion job's health and progress
- Watching a transformation job's health and progress
- Tracking whether a semantic model's refresh actually succeeded
- Setting up an alert to catch a failure early
- Tracking down why a pipeline run failed and fixing it
- Diagnosing why a dataflow run failed
- Debugging a notebook run that failed
- Troubleshooting a misbehaving Eventhouse
- Troubleshooting a misbehaving Eventstream
- Debugging a T-SQL statement that failed
- Fixing a broken or unreachable shortcut
- Speeding up a Lakehouse table with maintenance operations
- Making a slow pipeline run faster
- Tuning a Fabric warehouse for faster queries
- Improving throughput on real-time streaming components
- Tuning a Spark job to run faster and cheaper
- Making a slow query run faster
Coverage checked against the published exam guide on Aug 11, 2026.
These are independent practice questions, written against this certification's published exam guide. They are not the certification vendor's own questions, and not the real exam.