Securing data at the OneLake storage layer
OneLake security lets you define centralized, fine-grained access control (at the folder, table, row, and column level) directly on data stored in OneLake, primarily for Lakehouse items, so permissions are enforced consistently no matter which engine or tool queries the data. Learners must know how to create OneLake data access roles, scope them to specific folders/tables (with optional row/column filters), and assign them to users or groups. It's important to understand how this data-level security layers on top of, rather than replaces, workspace and item permissions.
Must-know
- OneLake security is configured through 'data access roles' defined on a Lakehouse item, letting you grant access to specific folders (under Files or Tables) or specific tables instead of the entire item.
- Row-level and column-level security can be added to a data access role for Delta tables by specifying filter predicates for rows and choosing which columns to include or exclude.
- Access enforced by OneLake data access roles applies uniformly across engines that read through OneLake, such as the SQL analytics endpoint and Power BI/Direct Lake, without needing separate per-engine configuration.
- If a user belongs to multiple OneLake data access roles, their effective access is the union of all permissions granted by those roles (most permissive combination wins).
- OneLake security governs data-level access only; a user still needs sufficient workspace role or item permission to open/use the Lakehouse itself before OneLake role permissions take effect.
- Lakehouse items ship with a default 'Read all' data access role granting full read access to users who already have item read permission, and admins can create narrower custom roles to restrict this default behavior.
A lakehouse contains three top-level folders: Finance, HR, and Sales, each holding files and tables for its department. The data engineering team wants the HR security group to read only the contents of the HR folder through OneLake, while members of that group should not see the Finance or Sales folders at all, and they should not receive workspace-level Viewer access. What should the data engineer configure?
What you have tried across DP-700's objectives, not a readiness score.
Implement and manage an analytics solution
- Tuning a workspace's Spark compute defaults and pool sizing
- Grouping and governing workspaces with a Fabric domain
- Setting per-workspace defaults for OneLake storage
- Standing up an Airflow job runtime inside a workspace
- Connecting a workspace to a Git repository
- Managing schema changes with a database project
- Promoting Fabric items across environments with a deployment pipeline
- Granting and restricting access at the workspace level
- Locking down who can open a single Fabric item
- Layering row, column, object, and file-level security rules
- Hiding sensitive column values behind a dynamic mask
- Classifying Fabric items with a sensitivity label
- Marking a trusted item as promoted or certified
- Reading a Fabric audit log to see who did what
- Securing data at the OneLake storage layer
- Picking the right build tool among a dataflow, a pipeline, and a notebook
- Kicking off a job on a schedule or in response to an event
- Chaining notebooks and pipelines together with parameters and dynamic expressions
Ingest and transform data
- Deciding between a full reload and an incremental load
- Shaping source data ahead of a dimensional-model load
- Landing a continuous stream of data into storage
- Matching a workload to the right Fabric data store
- Picking a transformation tool from dataflows, notebooks, KQL, or T-SQL
- Linking to external data without copying it via a OneLake shortcut
- Keeping a source database continuously replicated into Fabric
- Moving data into Fabric with a data pipeline
- Writing transform logic in PySpark, SQL, or KQL
- Flattening related tables into one wide, denormalized shape
- Rolling records up with group-by aggregations
- Dealing with duplicate rows, gaps, and data that arrives late
- Selecting the right engine for a real-time workload
- Weighing storage-in-place against a linked shortcut for a Real-Time Intelligence table
- Weighing an accelerated shortcut against a standard one for query speed
- Routing and reshaping live events with an Eventstream
- Handling a continuous flow of records with Spark's structured streaming
- Querying and reshaping event data with KQL
- Aggregating a stream over sliding or tumbling time windows
Monitor and optimize an analytics solution
- Watching an ingestion job's health and progress
- Watching a transformation job's health and progress
- Tracking whether a semantic model's refresh actually succeeded
- Setting up an alert to catch a failure early
- Tracking down why a pipeline run failed and fixing it
- Diagnosing why a dataflow run failed
- Debugging a notebook run that failed
- Troubleshooting a misbehaving Eventhouse
- Troubleshooting a misbehaving Eventstream
- Debugging a T-SQL statement that failed
- Fixing a broken or unreachable shortcut
- Speeding up a Lakehouse table with maintenance operations
- Making a slow pipeline run faster
- Tuning a Fabric warehouse for faster queries
- Improving throughput on real-time streaming components
- Tuning a Spark job to run faster and cheaper
- Making a slow query run faster
Coverage checked against the published exam guide on Aug 11, 2026.
These are independent practice questions, written against this certification's published exam guide. They are not the certification vendor's own questions, and not the real exam.