Skip to content

Pulling data through JDBC, ODBC, or REST clients and scheduling the job

External systems (databases, APIs, SaaS apps) can be pulled into Databricks by running JDBC/ODBC or REST client code inside a notebook, writing results as files to cloud storage or directly as Delta tables in Unity Catalog. These notebooks are then wrapped in Lakeflow Jobs so the ingestion runs on a schedule or trigger without manual intervention.

Must-know

  • Notebooks can use standard Python/Scala JDBC drivers (e.g. via Spark's spark.read.jdbc/format('jdbc')) to connect to relational sources, or REST/HTTP client libraries (e.g. requests) to pull data from APIs.
  • Spark's JDBC reader supports partitioning options (partitionColumn, lowerBound, upperBound, numPartitions) to parallelize large table pulls and avoid single-connection bottlenecks.
  • Data landed via REST calls typically arrives as JSON/text in driver memory or a local file first, then must be converted to a Spark DataFrame and written out (e.g. as Delta) rather than assumed to be distributed automatically.
  • Writing directly into Unity Catalog–governed tables requires the cluster or SQL warehouse to have a Unity Catalog-enabled access mode, and the identity running the job needs appropriate USE CATALOG/USE SCHEMA/CREATE TABLE grants.
  • Credentials for JDBC connections or API tokens should be stored using Databricks secrets (secret scopes) rather than hardcoded in notebook cells.
  • Lakeflow Jobs orchestrate these ingestion notebooks with schedules, triggers, retries, and alerting, and can chain them as tasks alongside downstream transformation steps.
Check this objectiveFree · always available

A data engineer connects to an on-premises PostgreSQL database from a notebook using a JDBC URL with the username and password embedded as plain-text strings, then schedules the notebook as a Lakeflow Job that runs nightly. A security review flags the plain-text credentials. Which approach removes the finding while keeping the job functional?

Your objective map0 tried · 0 right · 33 untouched

What you have tried across Databricks DEA's objectives, not a readiness score.

Coverage checked against the published exam guide on Jul 26, 2026.

These are independent practice questions, written against this certification's published exam guide. They are not the certification vendor's own questions, and not the real exam.