Joining and combining DataFrames with the different join and union types
Spark SQL and DataFrame APIs support standard SQL join types plus union operations for combining datasets. Join type and key selection affect result cardinality and null handling, while broadcast joins optimize performance when one side is small.
Must-know
- Inner join returns only matching rows from both DataFrames; left (outer) join keeps all rows from the left side, filling unmatched right-side columns with null.
- Multiple join keys are specified by passing a list of column names or an AND-combined boolean expression, e.g. df1.join(df2, ['col1','col2']) or df1.join(df2, (df1.a==df2.a) & (df1.b==df2.b)).
- Broadcast join (broadcast() hint or automatic via spark.sql.autoBroadcastJoinThreshold) sends a small DataFrame to all executors to avoid a costly shuffle when joining with a much larger table.
- Cross join produces the Cartesian product of two DataFrames (every row paired with every row) and can explode row counts quickly, so it must be used deliberately, often via crossJoin() or an explicit CROSS JOIN.
- union() (or unionByName()) combines rows from two DataFrames with the same schema and does NOT remove duplicates, unlike SQL UNION which dedupes; use distinct() after union to mimic that behavior.
- unionByName() matches columns by name rather than position, which avoids silently misaligned data when column order differs between DataFrames, and can optionally allow missing columns with allowMissingColumns=True.
A table customers has columns customer_id and region, and a table subscriptions has columns customer_id, region, and plan. Some customers have moved regions over time, so a match should only occur when both customer_id AND region agree between the two tables. Which code correctly implements this requirement?
What you have tried across Databricks DEA's objectives, not a readiness score.
Databricks Intelligence Platform
Data Ingestion and Loading
- Batch, streaming, and incremental loading patterns, and where the data comes from
- Loading files from cloud storage into governed tables with COPY INTO
- Landing data with Auto Loader, and handling schema enforcement and evolution
- Setting up Lakeflow Connect to ingest from enterprise sources reliably
- Pulling data through JDBC, ODBC, or REST clients and scheduling the job
- Choosing the right ingestion method for a given volume, frequency, and governance need
- Bringing semi-structured and unstructured data into governed Delta tables
Data Transformation and Modeling
- Cleaning bronze data into silver tables with PySpark and SQL
- Joining and combining DataFrames with the different join and union types
- Reshaping columns, rows, and arrays in a table
- Deduplicating and aggregating DataFrames
- Tuning Spark's core parameters and measuring what changed
- Building Gold-layer views and tables for BI and analytics
- Validating Silver and Gold datasets for quality
Working with Lakeflow Jobs
Implementing CI/CD
- Branching, committing, and opening pull requests from inside the Databricks workspace
- Promoting one codebase across dev, test, and prod with bundle variables and overrides
- Packaging and deploying jobs and pipelines with Automation Bundles
- Validating and managing bundle deployments from the Databricks CLI
Troubleshooting, Monitoring, and Optimization
- Spotting performance trends in a job's run history
- Reading job status, task graphs, and failure rates to monitor pipeline health
- Diagnosing skew, shuffle, and spill from Spark UI stage metrics
- What Liquid Clustering and predictive optimization actually do
- Diagnosing cluster startup failures, library conflicts, and out-of-memory errors
Governance and Security
Coverage checked against the published exam guide on Jul 25, 2026.
These are independent practice questions, written against this certification's published exam guide. They are not the certification vendor's own questions, and not the real exam.