
Jun-2026 Google Professional-Data-Engineer Actual Questions and 100% Cover Real Exam Questions
Professional-Data-Engineer Free Exam Questions and Answers PDF Updated on Jun-2026
To pass the Google Professional-Data-Engineer exam, candidates must have a solid understanding of data engineering concepts and techniques, as well as practical experience working with the Google Cloud Platform. They must be able to design and implement data processing systems that are secure, scalable, and efficient, and have the ability to troubleshoot and optimize these systems as needed. Professional-Data-Engineer exam is challenging and comprehensive, but passing it can open up many career opportunities in data engineering, especially for those interested in working with Google Cloud Platform.
NEW QUESTION # 102
You stream order data by using a Dataflow pipeline, and write the aggregated result to Memorystore. You provisioned a Memorystore for Redis instance with Basic Tier. 4 GB capacity, which is used by 40 clients for read-only access. You are expecting the number of read-only clients to increase significantly to a few hundred and you need to be able to support the demand. You want to ensure that read and write access availability is not impacted, and any changes you make can be deployed quickly. What should you do?
- A. Create a new Memorystore for Redis instance with Standard Tier Set capacity to 4 GB and read replica to No read replicas (high availability only). Delete the old instance.
- B. Create multiple new Memorystore for Redis instances with Basic Tier (4 GB capacity) Modify the Dataflow pipeline and new clients to use all instances
- C. Create a new Memorystore for Redis instance with Standard Tier Set capacity to 5 GB and create multiple read replicas Delete the old instance.
- D. Create a new Memorystore for Memcached instance Set a minimum of three nodes, and memory per node to 4 GB. Modify the Dataflow pipeline and all clients to use the Memcached instance Delete the old instance.
Answer: C
Explanation:
The Basic Tier of Memorystore for Redis provides a standalone Redis instance that is not replicated and does not support read replicas. This means that it cannot scale horizontally to handle more read requests, and it does not provide high availability or automatic failover. If the number of read-only clients increases significantly, the Basic Tier instance may not be able to handle the demand and may impact the read and write access availability. Therefore, option A is not a good solution, as it would require creating multiple Basic Tier instances and modifying the Dataflow pipeline and the clients to distribute the load among them. This would increase the complexity and the management overhead of the solution.
The Standard Tier of Memorystore for Redis provides a highly available Redis instance that supports replication and read replicas. Replication ensures that the data is backed up in another zone and can fail over automatically in case of a primary node failure. Read replicas allow scaling the read throughput by adding up to five replicas to an instance and using them for read-only queries. The Standard Tier also supports in-transit encryption and maintenance windows. Therefore, option D is the best solution, as it would create a new Standard Tier instance with a higher capacity (5 GB) and multiple read replicas to handle the increased demand. The old instance can be deleted after migrating the data to the new instance.
Option B is not a good solution, as it would create a new Standard Tier instance with the same capacity (4 GB) and no read replicas. This would not improve the read throughput or the availability of the solution. Option C is not a good solution, as it would create a new Memorystore for Memcached instance, which is a different service that uses a different protocol and data model than Redis. This would require changing the code of the Dataflow pipeline and the clients to use the Memcached protocol and data structures, which would take more time and effort than migrating to a new Redis instance. References: Redis tier capabilities | Memorystore for Redis | Google Cloud, Pricing | Memorystore for Redis | Google Cloud, What is Memorystore? | Google Cloud Blog, Working with GCP Memorystore - Simple Talk - Redgate Software
NEW QUESTION # 103
You are designing storage for two relational tables that are part of a 10-TB database on Google Cloud. You want to support transactions that scale horizontally. You also want to optimize data for range queries on non- key columns. What should you do?
- A. Use Cloud Spanner for storage. Add secondary indexes to support query patterns.
- B. Use Cloud Spanner for storage. Use Cloud Dataflow to transform data to support query patterns.
- C. Use Cloud SQL for storage. Use Cloud Dataflow to transform data to support query patterns.
- D. Use Cloud SQL for storage. Add secondary indexes to support query patterns.
Answer: B
Explanation:
Explanation/Reference: https://cloud.google.com/solutions/data-lifecycle-cloud-platform
NEW QUESTION # 104
You're using Bigtable for a real-time application, and you have a heavy load that is a mix of read and writes. You've recently identified an additional use case and need to perform hourly an analytical job to calculate certain statistics across the whole database. You need to ensure both the reliability of your production application as well as the analytical workload.
What should you do?
- A. Export Bigtable dump to GCS and run your analytical job on top of the exported files.
- B. Add a second cluster to an existing instance with a single-cluster routing, use live-traffic app profile for your regular workload and batch-analytics profile for the analytics workload.
- C. Add a second cluster to an existing instance with a multi-cluster routing, use live-traffic app profile for your regular workload and batch-analytics profile for the analytics workload.
- D. Increase the size of your existing cluster twice and execute your analytics workload on your new resized cluster.
Answer: C
Explanation:
https://cloud.google.com/bigtable/docs/replication-settings#batch-vs-serve
NEW QUESTION # 105
Your business users need a way to clean and prepare data before using the data for analysis. Your business users are less technically savvy and prefer to work with graphical user interfaces to define their transformations. After the data has been transformed, the business users want to perform their analysis directly in a spreadsheet. You need to recommend a solution that they can use. What should you do?
- A. Use Dataflow to clean the data, and write the results to BigQuery. Analyze the data by using Connected Sheets.
- B. Use Dataprep to clean the data, and write the results to BigQuery Analyze the data by using Connected Sheets.
- C. Use Dataprep to clean the data, and write the results to BigQuery Analyze the data by using Looker Studio.
- D. Use Dataflow to clean the data, and write the results to BigQuery. Analyze the data by using Looker Studio.
Answer: B
Explanation:
For business users who are less technically savvy and prefer graphical user interfaces, Dataprep is an ideal tool for cleaning and preparing data, as it offers a user-friendly interface for defining data transformations without the need for coding. Once the data is cleaned and prepared, writing the results to BigQuery allows for the storage and management of large datasets. Analyzing the data using Connected Sheets enables business users to work within the familiar environment of a spreadsheet, leveraging the power of BigQuery directly within Google Sheets. This solution aligns with the needs of the users and follows Google's recommended practices for data cleaning, preparation, and analysis.
References:
Connected Sheets | Google Sheets | Google for Developers
Professional Data Engineer Certification Exam Guide | Learn - Google Cloud Engineer Data in Google Cloud | Google Cloud Skills Boost - Qwiklabs
NEW QUESTION # 106
You are building a new application that you need to collect data from in a scalable way. Data arrives continuously from the application throughout the day, and you expect to generate approximately 150 GB of JSON data per day by the end of the year. Your requirements are:
* Decoupling producer from consumer
* Space and cost-efficient storage of the raw ingested data, which is to be stored indefinitely
* Near real-time SQL query
* Maintain at least 2 years of historical data, which will be queried with SQL Which pipeline should you use to meet these requirements?
- A. Create an application that publishes events to Cloud Pub/Sub, and create Spark jobs on Cloud Dataproc to convert the JSON data to Avro format, stored on HDFS on Persistent Disk.
- B. Create an application that writes to a Cloud SQL database to store the data. Set up periodic exports of the database to write to Cloud Storage and load into BigQuery.
- C. Create an application that provides an API. Write a tool to poll the API and write data to Cloud Storage as gzipped JSON files.
- D. Create an application that publishes events to Cloud Pub/Sub, and create a Cloud Dataflow pipeline that transforms the JSON event payloads to Avro, writing the data to Cloud Storage and BigQuery.
Answer: C
NEW QUESTION # 107
You set up a streaming data insert into a Redis cluster via a Kafka cluster. Both clusters are running on Compute Engine instances. You need to encrypt data at rest with encryption keys that you can create, rotate, and destroy as needed. What should you do?
- A. Create a dedicated service account, and use encryption at rest to reference your data stored in your Compute Engine cluster instances as part of your API service calls.
- B. Create encryption keys locally. Upload your encryption keys to Cloud Key Management Service. Use those keys to encrypt your data in all of the Compute Engine cluster instances.
- C. Create encryption keys in Cloud Key Management Service. Reference those keys in your API service calls when accessing the data in your Compute Engine cluster instances.
- D. Create encryption keys in Cloud Key Management Service. Use those keys to encrypt your data in all of the Compute Engine cluster instances.
Answer: B
NEW QUESTION # 108
Your company's customer and order databases are often under heavy load. This makes performing analytics against them difficult without harming operations. The databases are in a MySQL cluster, with nightly backups taken using mysqldump. You want to perform analytics with minimal impact on operations.
What should you do?
- A. Mount the backups to Google Cloud SQL, and then process the data using Google Cloud Dataproc.
- B. Connect an on-premises Apache Hadoop cluster to MySQL and perform ETL.
- C. Use an ETL tool to load the data from MySQL into Google BigQuery.
- D. Add a node to the MySQL cluster and build an OLAP cube there.
Answer: B
NEW QUESTION # 109
You are implementing security best practices on your data pipeline. Currently, you are manually executing jobs as the Project Owner. You want to automate these jobs by taking nightly batch files containing non- public information from Google Cloud Storage, processing them with a Spark Scala job on a Google Cloud Dataproc cluster, and depositing the results into Google BigQuery.
How should you securely run this workload?
- A. Restrict the Google Cloud Storage bucket so only you can see the files
- B. Use a user account with the Project Viewer role on the Cloud Dataproc cluster to read the batch files and write to BigQuery
- C. Use a service account with the ability to read the batch files and to write to BigQuery
- D. Grant the Project Owner role to a service account, and run the job with it
Answer: C
NEW QUESTION # 110
Which row keys are likely to cause a disproportionate number of reads and/or writes on a particular node in a Bigtable cluster (select 2 answers)?
- A. A timestamp followed by a stock symbol
- B. A stock symbol followed by a timestamp
- C. A sequential numeric ID
- D. A non-sequential numeric ID
Answer: A,C
Explanation:
...using a timestamp as the first element of a row key can cause a variety of problems. In brief, when a row key for a time series includes a timestamp, all of your writes will target a single node; fill that node; and then move onto the next node in the cluster, resulting in hotspotting. Suppose your system assigns a numeric ID to each of your application's users. You might be tempted to use the user's numeric ID as the row key for your table. However, since new users are more likely to be active users, this approach is likely to push most of your traffic to a small number of nodes. [https://cloud.google.com/bigtable/docs/schema- design] Reference: https://cloud.google.com/bigtable/docs/schema-design-time- series#ensure_that_your_row_key_avoids_hotspotting
NEW QUESTION # 111
You have a query that filters a BigQuery table using a WHERE clause on timestamp and ID columns. By using bq query - -dry_run you learn that the query triggers a full scan of the table, even though the filter on timestamp and ID select a tiny fraction of the overall data. You want to reduce the amount of data scanned by BigQuery with minimal changes to existing SQL queries. What should you do?
- A. Recreate the table with a partitioning column and clustering column.
- B. Use the bq query - -maximum_bytes_billed flag to restrict the number of bytes billed.
- C. Use the LIMIT keyword to reduce the number of rows returned.
- D. Create a separate table for each ID.
Answer: C
NEW QUESTION # 112
Google Cloud Bigtable indexes a single value in each row. This value is called the _______.
- A. master key
- B. unique key
- C. row key
- D. primary key
Answer: C
Explanation:
Cloud Bigtable is a sparsely populated table that can scale to billions of rows and thousands of columns, allowing you to store terabytes or even petabytes of data. A single value in each row is indexed; this value is known as the row key.
Reference: https://cloud.google.com/bigtable/docs/overview
NEW QUESTION # 113
Your neural network model is taking days to train. You want to increase the training speed. What can you do?
- A. Subsample your training dataset.
- B. Subsample your test dataset.
- C. Increase the number of layers in your neural network.
- D. Increase the number of input features to your model.
Answer: A
Explanation:
Subsampling is the method to increase the training speed.
NEW QUESTION # 114
You work for an advertising company, and you've developed a Spark ML model to predict click-through rates at advertisement blocks. You've been developing everything at your on-premises data center, and now your company is migrating to Google Cloud. Your data center will be migrated to BigQuery. You periodically retrain your Spark ML models, so you need to migrate existing training pipelines to Google Cloud. What should you do?
- A. Use Cloud Dataproc for training existing Spark ML models, but start reading data directly from BigQuery
- B. Rewrite your models on TensorFlow, and start using Cloud ML Engine
- C. Spin up a Spark cluster on Compute Engine, and train Spark ML models on the data exported from BigQuery
- D. Use Cloud ML Engine for training existing Spark ML models
Answer: A
Explanation:
https://cloud.google.com/dataproc/docs/tutorials/bigquery-sparkml
NEW QUESTION # 115
Which of the following statements is NOT true regarding Bigtable access roles?
- A. Using IAM roles, you cannot give a user access to only one table in a project, rather than all tables in a project.
- B. To give a user access to only one table in a project, you must configure access through your application.
- C. To give a user access to only one table in a project, grant the user the Bigtable Editor role for that table.
- D. You can configure access control only at the project level.
Answer: C
Explanation:
For Cloud Bigtable, you can configure access control at the project level. For example, you can grant the ability to:
Read from, but not write to, any table within the project.
Read from and write to any table within the project, but not manage instances.
Read from and write to any table within the project, and manage instances.
NEW QUESTION # 116
You operate a logistics company, and you want to improve event delivery reliability for vehicle-based sensors. You operate small data centers around the world to capture these events, but leased lines that provide connectivity from your event collection infrastructure to your event processing infrastructure are unreliable, with unpredictable latency. You want to address this issue in the most cost-effective way. What should you do?
- A. Establish a Cloud Interconnect between all remote data centers and Google.
- B. Deploy small Kafka clusters in your data centers to buffer events.
- C. Write a Cloud Dataflow pipeline that aggregates all data in session windows.
- D. Have the data acquisition devices publish data to Cloud Pub/Sub.
Answer: B
NEW QUESTION # 117
Your company uses a proprietary system to send inventory data every 6 hours to a data ingestion service in the cloud. Transmitted data includes a payload of several fields and the timestamp of the transmission. If there are any concerns about a transmission, the system re-transmits the data. How should you deduplicate the data most efficiency?
- A. Store each data entry as the primary key in a separate database and apply an index.
- B. Compute the hash value of each data entry, and compare it with all historical data.
- C. Assign global unique identifiers (GUID) to each data entry.
- D. Maintain a database table to store the hash value and other metadata for each data entry.
Answer: D
Explanation:
Using Hash values we can remove duplicate values from a database. Hashvalues will be same for duplicate data and thus can be easily rejected.
NEW QUESTION # 118
You designed a database for patient records as a pilot project to cover a few hundred patients in three clinics.
Your design used a single database table to represent all patients and their visits, and you used self-joins to generate reports. The server resource utilization was at 50%. Since then, the scope of the project has expanded. The database must now store 100 times more patient records. You can no longer run the reports, because they either take too long or they encounter errors with insufficient compute resources. How should you adjust the database design?
- A. Normalize the master patient-record table into the patient table and the visits table, and create other necessary tables to avoid self-join.
- B. Partition the table into smaller tables, with one for each clinic. Run queries against the smaller table pairs, and use unions for consolidated reports.
- C. Shard the tables into smaller ones based on date ranges, and only generate reports with prespecified date ranges.
- D. Add capacity (memory and disk space) to the database server by the order of 200.
Answer: A
Explanation:
Explanation
NEW QUESTION # 119
Your new customer has requested daily reports that show their net consumption of Google Cloud compute resources and who used the resources. You need to quickly and efficiently generate these daily reports. What should you do?
- A. Do daily exports of Cloud Logging data to BigQuery. Create views filtering by project, log type, resource, and user.
- B. Export Cloud Logging data to Cloud Storage in CSV format. Cleanse the data using Dataprep, filtering by project, resource, and user.
- C. Filter data in Cloud Logging by project, resource, and user; then export the data in CSV format.
- D. Filter data in Cloud Logging by project, log type, resource, and user, then import the data into BigQuery.
Answer: C
Explanation:
https://cloud.google.com/logging/docs/view/logs-explorer-interface?cloudshell=true
NEW QUESTION # 120
Your company is currently setting up data pipelines for their campaign. For all the Google Cloud Pub/Sub
streaming data, one of the important business requirements is to be able to periodically identify the inputs
and their timings during their campaign. Engineers have decided to use windowing and transformation in
Google Cloud Dataflow for this purpose. However, when testing this feature, they find that the Cloud
Dataflow job fails for the all streaming insert. What is the most likely cause of this problem?
- A. They have not set the triggers to accommodate the data coming in late, which causes the job to fail
- B. They have not applied a non-global windowing function, which causes the job to fail when the pipeline
is created - C. They have not assigned the timestamp, which causes the job to fail
- D. They have not applied a global windowing function, which causes the job to fail when the pipeline is
created
Answer: D
NEW QUESTION # 121
......
Google Professional-Data-Engineer certification exam covers a broad range of topics, including data processing systems, data modeling, data analysis, data visualization, and machine learning. It requires a strong understanding of Google Cloud Platform products and services, such as BigQuery, Dataflow, Dataproc, and Pub/Sub. Professional-Data-Engineer exam also tests the ability to design and implement solutions that are scalable, efficient, and secure.
Google Professional-Data-Engineer Real 2026 Braindumps Mock Exam Dumps: https://www.exam4labs.com/Professional-Data-Engineer-practice-torrent.html
Latest Professional-Data-Engineer Exam Dumps Recently Updated 403 Questions: https://drive.google.com/open?id=13ckatUFSP-Q3WM4KPeFHA5lwROZ7Wxul