Skip to content

Telling CSV, JSON, Parquet, Avro, and relational tables apart, and where each fits

As a Data Practitioner, you must recognize common data formats and know how Google Cloud services handle them, since the format determines ingestion method, schema handling, and query performance. CSV and JSON are common text-based, self-describing (JSON) or delimited (CSV) formats used for smaller or human-readable data, while Avro and Parquet are binary formats optimized for big data pipelines. Structured database tables refer to relational data typically extracted via database migration or federated query tools like Datastream or BigQuery federated queries.

Must-know

  • CSV files are row-based, human-readable, and require explicit schema definition or autodetect when loading into BigQuery, and they do not support nested or repeated fields.
  • JSON (specifically newline-delimited JSON, NDJSON) supports nested and repeated fields natively and is commonly used for semi-structured data ingestion into BigQuery.
  • Apache Avro is a row-based binary format with embedded schema, making it efficient for write-heavy workloads and schema evolution, and is the fastest format for BigQuery load jobs since it doesn't require separate schema parsing.
  • Apache Parquet is a column-based binary format optimized for analytical (read-heavy) workloads, offering high compression and efficient columnar scans, making it ideal for BigQuery external tables and Dataproc/Spark analytics.
  • Structured database tables (e.g., from Cloud SQL, PostgreSQL, MySQL) are typically ingested into BigQuery via Datastream for change data capture or via federated queries, preserving relational schema and data types.
  • Choosing the right format matters for cost and performance: columnar formats (Parquet) reduce BigQuery scan costs for analytical queries, while row-based formats (Avro, CSV, JSON) are better suited for transactional or streaming ingestion.
Check this objectiveFree · always available

A team is building a streaming pipeline that publishes sensor readings to Pub/Sub. New sensor types are added regularly, so the message schema changes often, and consumers must be able to read old and new messages correctly. The team wants the schema to travel with each message and wants the serialized payload to stay compact using binary encoding. Which format should they choose for the messages?

Your objective map0 tried · 0 right · 41 untouched

What you have tried across GCP ADP's objectives, not a readiness score.

Coverage checked against the published exam guide on Aug 12, 2026.

These are independent practice questions, written against this certification's published exam guide. They are not the certification vendor's own questions, and not the real exam.