Getting files and tables loaded with a CLI, a transfer service, or a client library
The Associate Data Practitioner exam expects you to choose the right ingestion tool based on data source, frequency, and destination: CLI tools for ad hoc or scripted loads, Storage Transfer Service for bulk/scheduled object storage transfers, BigQuery Data Transfer Service for recurring loads from SaaS and other Google services, and client libraries for programmatic, application-driven ingestion. Matching the tool to the use case (one-time vs. recurring, source type, and required transformations) is a core exam skill.
Must-know
- gsutil or gcloud storage commands (e.g., gcloud storage cp) are used for manual or scripted uploads to Cloud Storage, while bq load handles command-line loads into BigQuery from local files or Cloud Storage.
- Storage Transfer Service is designed for large-scale, scheduled, or repeated transfers into Cloud Storage from sources like other cloud providers (S3, Azure Blob), on-premises data (via agents), or Cloud Storage-to-Cloud Storage moves.
- BigQuery Data Transfer Service automates recurring, managed ingestion directly into BigQuery from supported sources such as Google Ads, Google Ad Manager, YouTube, Google Merchant Center, Amazon S3, and Cloud Storage, on a schedule you define.
- Client libraries (available in Python, Java, Go, Node.js, etc.) are the correct choice when ingestion must be embedded in custom application code, requires fine-grained control, or needs to integrate with streaming/real-time pipelines.
- For streaming or event-driven ingestion into BigQuery, the exam expects awareness that the Storage Write API (via client libraries) is the current recommended method, not older tabledata.insertAll patterns.
- A key exam gotcha: bq load and BigQuery Data Transfer Service can only load data already in supported formats/locations, whereas Storage Transfer Service moves raw objects but does not parse or load them into BigQuery tables. Know which tool stops at storage and which completes the load into BigQuery.
A data engineering team needs to migrate roughly 40 TB of log files from an Amazon S3 bucket to a Cloud Storage bucket on a recurring weekly basis, with minimal ongoing operational effort. Which approach should they use?
What you have tried across GCP ADP's objectives, not a readiness score.
Data Preparation and Ingestion
- When to load first and when to transform first, and what sits between the two
- Picking a way to move existing data into Google Cloud
- Judging whether a dataset is trustworthy enough to build on
- Fixing messy records before they reach a report
- Telling CSV, JSON, Parquet, Avro, and relational tables apart, and where each fits
- Picking how to pull data out of a source system
- Matching a workload to the right storage or database service
- Getting files and tables loaded with a CLI, a transfer service, or a client library
Data Analysis and Presentation
- Writing BigQuery SQL that answers a reporting question
- Exploring and charting data inside a hosted notebook
- Turning a question from the business into an analysis that settles it
- Building a dashboard and getting it in front of the right people
- Deciding whether a job calls for Looker or for Looker Studio
- Editing LookML to change what a model exposes
- Spotting a problem worth solving with BigQuery ML or AutoML
- Calling a hosted Google language model straight from BigQuery
- Sequencing a machine learning project from raw data to served predictions
- Building, fitting, and scoring a model with SQL alone
- Running predictions against a model you already trained
- Keeping trained models catalogued in one place
Data Pipeline Orchestration
- Matching a transformation job to Dataproc, Dataflow, Dataform, or a managed alternative
- Weighing whether the transform belongs before or after the load
- Assembling the services a simple transformation pipeline needs
- Putting a query on a schedule and keeping it running
- Watching a Dataflow job and spotting where it stalls
- Reading logs and metrics to work out what a pipeline actually did
- Choosing what should drive a multi-step workflow
- Streaming messages into BigQuery as they arrive rather than in batches
- Wiring a trigger so one event starts the next step
Data Management
- Granting only the access a person or service actually needs
- Controlling who can read a bucket, and what uniform access changes
- Sharing a dataset with another team or company without copying it
- Matching a storage class to how often the data gets read
- Expiring old data automatically so it stops costing money
- Picking somewhere to park data that must be kept but is rarely read
- Comparing the managed backup and restore options across services
- Working out when a second copy is worth what it costs
- Regions, dual-regions, multi-regions, and zones as redundancy choices
- Deciding who should hold the encryption keys
- What a key management service does for creating, rotating, and revoking keys
- Protecting data on the wire versus data sitting on a disk
Coverage checked against the published exam guide on Aug 12, 2026.
These are independent practice questions, written against this certification's published exam guide. They are not the certification vendor's own questions, and not the real exam.