Managed versus external tables in Unity Catalog, and converting between them
Unity Catalog supports managed and external tables, differing in who owns the underlying data lifecycle. Managed tables store data in the Unity Catalog managed storage location and are dropped along with their data, while external tables reference data at an existing cloud storage path via a location, and dropping them only removes the metadata.
Must-know
- Managed tables store data under the metastore/catalog/schema managed storage location; DROP TABLE deletes both metadata and underlying data files.
- External tables require an external location and storage credential registered in Unity Catalog, and DROP TABLE removes only the table metadata, leaving data files intact.
- CREATE TABLE without a LOCATION clause creates a managed table; specifying LOCATION 'path' creates an external table.
- You cannot directly convert a managed table to external (or vice versa) in place; the common approach is CREATE TABLE ... AS SELECT (CTAS) or CREATE TABLE ... LOCATION to copy/rebuild data into the desired table type.
- ALTER TABLE can rename a table or change properties/comments and column definitions, but ownership, location type, and storage credential associations are set at creation and not simply toggled via ALTER.
- Managed tables benefit fully from Unity Catalog governance features like Predictive Optimization and automatic file layout optimization, whereas external tables are typically used for data shared across platforms or requiring specific storage locations/lifecycle control outside Databricks.
A data engineer runs DROP TABLE sales.transactions and the table was a managed table in Unity Catalog. What happens to the underlying data files?
What you have tried across Databricks DEA's objectives, not a readiness score.
Databricks Intelligence Platform
Data Ingestion and Loading
- Batch, streaming, and incremental loading patterns, and where the data comes from
- Loading files from cloud storage into governed tables with COPY INTO
- Landing data with Auto Loader, and handling schema enforcement and evolution
- Setting up Lakeflow Connect to ingest from enterprise sources reliably
- Pulling data through JDBC, ODBC, or REST clients and scheduling the job
- Choosing the right ingestion method for a given volume, frequency, and governance need
- Bringing semi-structured and unstructured data into governed Delta tables
Data Transformation and Modeling
- Cleaning bronze data into silver tables with PySpark and SQL
- Joining and combining DataFrames with the different join and union types
- Reshaping columns, rows, and arrays in a table
- Deduplicating and aggregating DataFrames
- Tuning Spark's core parameters and measuring what changed
- Building Gold-layer views and tables for BI and analytics
- Validating Silver and Gold datasets for quality
Working with Lakeflow Jobs
Implementing CI/CD
- Branching, committing, and opening pull requests from inside the Databricks workspace
- Promoting one codebase across dev, test, and prod with bundle variables and overrides
- Packaging and deploying jobs and pipelines with Automation Bundles
- Validating and managing bundle deployments from the Databricks CLI
Troubleshooting, Monitoring, and Optimization
- Spotting performance trends in a job's run history
- Reading job status, task graphs, and failure rates to monitor pipeline health
- Diagnosing skew, shuffle, and spill from Spark UI stage metrics
- What Liquid Clustering and predictive optimization actually do
- Diagnosing cluster startup failures, library conflicts, and out-of-memory errors
Governance and Security
Coverage checked against the published exam guide on Jul 26, 2026.
These are independent practice questions, written against this certification's published exam guide. They are not the certification vendor's own questions, and not the real exam.