Spotting a problem worth solving with BigQuery ML or AutoML
BigQuery ML lets you build and deploy machine learning models directly on data stored in BigQuery using familiar SQL syntax, ideal for teams with SQL skills who want to avoid data movement. AutoML (via Vertex AI) provides a low-code/no-code interface for training high-quality custom models on structured, image, text, and video data when more advanced model types or greater customization are needed. Choosing between them depends on data location, team skillset, model complexity, and whether the use case fits BigQuery ML's supported algorithms.
Must-know
- BigQuery ML is best for structured/tabular data already in BigQuery and supports algorithms like linear/logistic regression, k-means clustering, matrix factorization, time-series forecasting (
ARIMA_PLUS), boosted trees (XGBoost), and deep neural networks, plus the ability to import TensorFlow models. - AutoML (part of Vertex AI) is best when you need to train models on unstructured data (images, text, video) or want Google's automated architecture search and hyperparameter tuning for higher accuracy without writing model code.
- BigQuery ML keeps data and compute in BigQuery, eliminating ETL/export steps and reducing latency for training and batch prediction on large datasets.
- Use BigQuery ML for quick iteration and embedding ML directly into analytics/BI workflows via SQL; use AutoML/Vertex AI when you need advanced model customization, explainability, online prediction endpoints, or MLOps pipeline integration.
- BigQuery ML also supports remote inference by calling pretrained Vertex AI models (e.g., for text generation or embeddings) directly from SQL, blurring the line between the two for certain use cases.
- For real-time/low-latency online predictions or complex custom architectures beyond BigQuery ML's supported model types, Vertex AI (custom training or AutoML) is the appropriate choice rather than BigQuery ML, which is optimized primarily for batch prediction.
A retail data analyst has two years of transaction history stored in BigQuery tables. The analyst wants to build a model that predicts which customers are likely to churn next month, using only SQL and without moving the data out of BigQuery. Which approach best meets these requirements?
What you have tried across GCP ADP's objectives, not a readiness score.
Data Preparation and Ingestion
- When to load first and when to transform first, and what sits between the two
- Picking a way to move existing data into Google Cloud
- Judging whether a dataset is trustworthy enough to build on
- Fixing messy records before they reach a report
- Telling CSV, JSON, Parquet, Avro, and relational tables apart, and where each fits
- Picking how to pull data out of a source system
- Matching a workload to the right storage or database service
- Getting files and tables loaded with a CLI, a transfer service, or a client library
Data Analysis and Presentation
- Writing BigQuery SQL that answers a reporting question
- Exploring and charting data inside a hosted notebook
- Turning a question from the business into an analysis that settles it
- Building a dashboard and getting it in front of the right people
- Deciding whether a job calls for Looker or for Looker Studio
- Editing LookML to change what a model exposes
- Spotting a problem worth solving with BigQuery ML or AutoML
- Calling a hosted Google language model straight from BigQuery
- Sequencing a machine learning project from raw data to served predictions
- Building, fitting, and scoring a model with SQL alone
- Running predictions against a model you already trained
- Keeping trained models catalogued in one place
Data Pipeline Orchestration
- Matching a transformation job to Dataproc, Dataflow, Dataform, or a managed alternative
- Weighing whether the transform belongs before or after the load
- Assembling the services a simple transformation pipeline needs
- Putting a query on a schedule and keeping it running
- Watching a Dataflow job and spotting where it stalls
- Reading logs and metrics to work out what a pipeline actually did
- Choosing what should drive a multi-step workflow
- Streaming messages into BigQuery as they arrive rather than in batches
- Wiring a trigger so one event starts the next step
Data Management
- Granting only the access a person or service actually needs
- Controlling who can read a bucket, and what uniform access changes
- Sharing a dataset with another team or company without copying it
- Matching a storage class to how often the data gets read
- Expiring old data automatically so it stops costing money
- Picking somewhere to park data that must be kept but is rarely read
- Comparing the managed backup and restore options across services
- Working out when a second copy is worth what it costs
- Regions, dual-regions, multi-regions, and zones as redundancy choices
- Deciding who should hold the encryption keys
- What a key management service does for creating, rotating, and revoking keys
- Protecting data on the wire versus data sitting on a disk
Coverage checked against the published exam guide on Aug 12, 2026.
These are independent practice questions, written against this certification's published exam guide. They are not the certification vendor's own questions, and not the real exam.