Telling CSV, JSON, Parquet, Avro, and relational tables apart, and where each fits
As a Data Practitioner, you must recognize common data formats and know how Google Cloud services handle them, since the format determines ingestion method, schema handling, and query performance. CSV and JSON are common text-based, self-describing (JSON) or delimited (CSV) formats used for smaller or human-readable data, while Avro and Parquet are binary formats optimized for big data pipelines. Structured database tables refer to relational data typically extracted via database migration or federated query tools like Datastream or BigQuery federated queries.
Must-know
- CSV files are row-based, human-readable, and require explicit schema definition or autodetect when loading into BigQuery, and they do not support nested or repeated fields.
- JSON (specifically newline-delimited JSON, NDJSON) supports nested and repeated fields natively and is commonly used for semi-structured data ingestion into BigQuery.
- Apache Avro is a row-based binary format with embedded schema, making it efficient for write-heavy workloads and schema evolution, and is the fastest format for BigQuery load jobs since it doesn't require separate schema parsing.
- Apache Parquet is a column-based binary format optimized for analytical (read-heavy) workloads, offering high compression and efficient columnar scans, making it ideal for BigQuery external tables and Dataproc/Spark analytics.
- Structured database tables (e.g., from Cloud SQL, PostgreSQL, MySQL) are typically ingested into BigQuery via Datastream for change data capture or via federated queries, preserving relational schema and data types.
- Choosing the right format matters for cost and performance: columnar formats (Parquet) reduce BigQuery scan costs for analytical queries, while row-based formats (Avro, CSV, JSON) are better suited for transactional or streaming ingestion.
A team is building a streaming pipeline that publishes sensor readings to Pub/Sub. New sensor types are added regularly, so the message schema changes often, and consumers must be able to read old and new messages correctly. The team wants the schema to travel with each message and wants the serialized payload to stay compact using binary encoding. Which format should they choose for the messages?
What you have tried across GCP ADP's objectives, not a readiness score.
Data Preparation and Ingestion
- When to load first and when to transform first, and what sits between the two
- Picking a way to move existing data into Google Cloud
- Judging whether a dataset is trustworthy enough to build on
- Fixing messy records before they reach a report
- Telling CSV, JSON, Parquet, Avro, and relational tables apart, and where each fits
- Picking how to pull data out of a source system
- Matching a workload to the right storage or database service
- Getting files and tables loaded with a CLI, a transfer service, or a client library
Data Analysis and Presentation
- Writing BigQuery SQL that answers a reporting question
- Exploring and charting data inside a hosted notebook
- Turning a question from the business into an analysis that settles it
- Building a dashboard and getting it in front of the right people
- Deciding whether a job calls for Looker or for Looker Studio
- Editing LookML to change what a model exposes
- Spotting a problem worth solving with BigQuery ML or AutoML
- Calling a hosted Google language model straight from BigQuery
- Sequencing a machine learning project from raw data to served predictions
- Building, fitting, and scoring a model with SQL alone
- Running predictions against a model you already trained
- Keeping trained models catalogued in one place
Data Pipeline Orchestration
- Matching a transformation job to Dataproc, Dataflow, Dataform, or a managed alternative
- Weighing whether the transform belongs before or after the load
- Assembling the services a simple transformation pipeline needs
- Putting a query on a schedule and keeping it running
- Watching a Dataflow job and spotting where it stalls
- Reading logs and metrics to work out what a pipeline actually did
- Choosing what should drive a multi-step workflow
- Streaming messages into BigQuery as they arrive rather than in batches
- Wiring a trigger so one event starts the next step
Data Management
- Granting only the access a person or service actually needs
- Controlling who can read a bucket, and what uniform access changes
- Sharing a dataset with another team or company without copying it
- Matching a storage class to how often the data gets read
- Expiring old data automatically so it stops costing money
- Picking somewhere to park data that must be kept but is rarely read
- Comparing the managed backup and restore options across services
- Working out when a second copy is worth what it costs
- Regions, dual-regions, multi-regions, and zones as redundancy choices
- Deciding who should hold the encryption keys
- What a key management service does for creating, rotating, and revoking keys
- Protecting data on the wire versus data sitting on a disk
Coverage checked against the published exam guide on Aug 12, 2026.
These are independent practice questions, written against this certification's published exam guide. They are not the certification vendor's own questions, and not the real exam.