Pulling data through JDBC, ODBC, or REST clients and scheduling the job
External systems (databases, APIs, SaaS apps) can be pulled into Databricks by running JDBC/ODBC or REST client code inside a notebook, writing results as files to cloud storage or directly as Delta tables in Unity Catalog. These notebooks are then wrapped in Lakeflow Jobs so the ingestion runs on a schedule or trigger without manual intervention.
1 · Learn the must-know
- Notebooks can use standard Python/Scala JDBC drivers (e.g. via Spark's spark.read.jdbc/format('jdbc')) to connect to relational sources, or REST/HTTP client libraries (e.g. requests) to pull data from APIs.
- Spark's JDBC reader supports partitioning options (partitionColumn, lowerBound, upperBound, numPartitions) to parallelize large table pulls and avoid single-connection bottlenecks.
- Data landed via REST calls typically arrives as JSON/text in driver memory or a local file first, then must be converted to a Spark DataFrame and written out (e.g. as Delta) rather than assumed to be distributed automatically.
- Writing directly into Unity Catalog–governed tables requires the cluster or SQL warehouse to have a Unity Catalog-enabled access mode, and the identity running the job needs appropriate USE CATALOG/USE SCHEMA/CREATE TABLE grants.
- Credentials for JDBC connections or API tokens should be stored using Databricks secrets (secret scopes) rather than hardcoded in notebook cells.
- Lakeflow Jobs orchestrate these ingestion notebooks with schedules, triggers, retries, and alerting, and can chain them as tasks alongside downstream transformation steps.
2 · Check your understanding
A data engineer connects to an on-premises PostgreSQL database from a notebook using a JDBC URL with the username and password embedded as plain-text strings, then schedules the notebook as a Lakeflow Job that runs nightly. A security review flags the plain-text credentials. Which approach removes the finding while keeping the job functional?
What you have tried across Databricks DEA's objectives, not a readiness score.
Databricks Intelligence Platform6% of the exam0 of 2 tried
Data Ingestion and Loading21% of the exam0 of 7 tried
Data Transformation and Modeling22% of the exam0 of 7 tried
Working with Lakeflow Jobs16% of the exam0 of 4 tried
Implementing CI/CD10% of the exam0 of 4 tried
Troubleshooting, Monitoring, and Optimization10% of the exam0 of 5 tried
Governance and Security15% of the exam0 of 4 tried
3 · Keep going
Ready for more? Take a weighted mock or try free practice questions.