Professional-Data-Engineer Questions Prepare with Learning Information! 2024 Regularly updated [Q158-Q176]

Share

Professional-Data-Engineer Questions Prepare with Learning Information! 2024 Regularly updated

Get Professional-Data-Engineer Products Practice Material for Professional-Data-Engineer Exam Question Preparation


To become a Google Professional-Data-Engineer, candidates need to pass the certification exam. Professional-Data-Engineer exam consists of multiple-choice and scenario-based questions that assess a candidate's ability to design, build, and manage data processing systems on the Google Cloud Platform. Professional-Data-Engineer exam can be taken online or in-person at a proctored testing center. Candidates have two hours to complete the exam, and they must score at least 70% to pass.

 

NEW QUESTION # 158
Which Java SDK class can you use to run your Dataflow programs locally?

  • A. MachineRunner
  • B. DirectPipelineRunner
  • C. LocalPipelineRunner
  • D. LocalRunner

Answer: B

Explanation:
DirectPipelineRunner allows you to execute operations in the pipeline directly, without any optimization. Useful for small local execution and tests Reference: https://cloud.google.com/dataflow/java- sdk/JavaDoc/com/google/cloud/dataflow/sdk/runners/DirectPipelineRunner


NEW QUESTION # 159
Your company is in a highly regulated industry. One of your requirements is to ensure individual users
have access only to the minimum amount of information required to do their jobs. You want to enforce this
requirement with Google BigQuery. Which three approaches can you take? (Choose three.)

  • A. Segregate data across multiple tables or databases.
  • B. Restrict access to tables by role.
  • C. Disable writes to certain tables.
  • D. Restrict BigQuery API access to approved users.
  • E. Use Google Stackdriver Audit Logging to determine policy violations.
  • F. Ensure that the data is encrypted at all times.

Answer: B,D,E


NEW QUESTION # 160
Your company has hired a new data scientist who wants to perform complicated analyses across very large datasets stored in Google Cloud Storage and in a Cassandra cluster on Google Compute Engine. The scientist primarily wants to create labelled data sets for machine learning projects, along with some visualization tasks.
She reports that her laptop is not powerful enough to perform her tasks and it is slowing her down. You want to help her perform her tasks. What should you do?

  • A. Deploy Google Cloud Datalab to a virtual machine (VM) on Google Compute Engine.
  • B. Run a local version of Jupiter on the laptop.
  • C. Grant the user access to Google Cloud Shell.
  • D. Host a visualization tool on a VM on Google Compute Engine.

Answer: C


NEW QUESTION # 161
You have enabled the free integration between Firebase Analytics and Google BigQuery. Firebase now
automatically creates a new table daily in BigQuery in the format app_events_YYYYMMDD.You want to
query all of the tables for the past 30 days in legacy SQL. What should you do?

  • A. Use SELECT IF.(date >= YYYY-MM-DD AND date <= YYYY-MM-DD
  • B. Use the WHERE_PARTITIONTIMEpseudo column
  • C. Use WHEREdate BETWEEN YYYY-MM-DD AND YYYY-MM-DD
  • D. Use the TABLE_DATE_RANGEfunction

Answer: D

Explanation:
Explanation/Reference:
Reference: https://cloud.google.com/blog/products/gcp/using-bigquery-and-firebase-analytics-to-
understand-your-mobile-app?hl=am


NEW QUESTION # 162
Which software libraries are supported by Cloud Machine Learning Engine?

  • A. Theano and TensorFlow
  • B. TensorFlow
  • C. Theano and Torch
  • D. TensorFlow and Torch

Answer: B

Explanation:
Cloud ML Engine mainly does two things:
Enables you to train machine learning models at scale by running TensorFlow training applications in the cloud.
Hosts those trained models for you in the cloud so that you can use them to get predictions about new data.
Reference: https://cloud.google.com/ml-engine/docs/technical-overview#what_it_does


NEW QUESTION # 163
You need to store and analyze social media postings in Google BigQuery at a rate of 10,000 messages per minute in near real-time. Initially, design the application to use streaming inserts for individual postings. Your application also performs data aggregations right after the streaming inserts. You discover that the queries after streaming inserts do not exhibit strong consistency, and reports from the queries might miss in-flight dat
a. How can you adjust your application design?

  • A. Load the original message to Google Cloud SQL, and export the table every hour to BigQuery via streaming inserts.
  • B. Estimate the average latency for data availability after streaming inserts, and always run queries after waiting twice as long.
  • C. Convert the streaming insert code to batch load for individual messages.
  • D. Re-write the application to load accumulated data every 2 minutes.

Answer: B

Explanation:
The data is first comes to buffer and then written to Storage. If we are running queries in buffer we will face above mentioned issues. If we wait for the bigquery to write the data to storage then we won't face the issue. So We need to wait till it's written tio storage


NEW QUESTION # 164
You create an important report for your large team in Google Data Studio 360. The report uses Google BigQuery as its data source. You notice that visualizations are not showing data that is less than 1 hour old. What should you do?

  • A. Disable caching by editing the report settings.
  • B. Refresh your browser tab showing the visualizations.
  • C. Clear your browser history for the past hour then reload the tab showing the virtualizations.
  • D. Disable caching in BigQuery by editing table details.

Answer: A

Explanation:
Reference https://support.google.com/datastudio/answer/7020039?hl=en


NEW QUESTION # 165
An organization maintains a Google BigQuery dataset that contains tables with user-level dat A.
They want to expose aggregates of this data to other Google Cloud projects, while still controlling access to the user-level data. Additionally, they need to minimize their overall storage cost and ensure the analysis cost for other projects is assigned to those projects. What should they do?

  • A. Create and share a new dataset and table that contains the aggregate results.
  • B. Create and share an authorized view that provides the aggregate results.
  • C. Create and share a new dataset and view that provides the aggregate results.
  • D. Create dataViewer Identity and Access Management (IAM) roles on the dataset to enable sharing.

Answer: D


NEW QUESTION # 166
You work for a large fast food restaurant chain with over 400,000 employees. You store employee information in Google BigQuery in a Users table consisting of a FirstName field and a LastName field. A member of IT is building an application and asks you to modify the schema and data in BigQuery so the application can query a FullName field consisting of the value of the FirstName field concatenated with a space, followed by the value of the LastName field for each employee. How can you make that data available while minimizing cost?

  • A. Add a new column called FullName to the Users table. Run an UPDATE statement that updates the FullName column for each user with the concatenation of the FirstName and LastName values.
  • B. Use BigQuery to export the data for the table to a CSV file. Create a Google Cloud Dataproc job to process the CSV file and output a new CSV file containing the proper values for FirstName, LastName and FullName. Run a BigQuery load job to load the new CSV file into BigQuery.
  • C. Create a Google Cloud Dataflow job that queries BigQuery for the entire Users table, concatenates the FirstName value and LastName value for each user, and loads the proper values for FirstName, LastName, and FullName into a new table in BigQuery.
  • D. Create a view in BigQuery that concatenates the FirstName and LastName field values to produce the FullName.

Answer: B

Explanation:
Import and Export to Bigquery from Cloud Storage is FREE. Also, when u store the csv files, Cloud Storage is cheaper than Bigquery. For processing Dataproc is cheaper than Dataflow.


NEW QUESTION # 167
Which of the following job types are supported by Cloud Dataproc (select 3 answers)?

  • A. Pig
  • B. Hive
  • C. Spark
  • D. YARN

Answer: A,B,C

Explanation:
Explanation
Cloud Dataproc provides out-of-the box and end-to-end support for many of the most popular job types, including Spark, Spark SQL, PySpark, MapReduce, Hive, and Pig jobs.
Reference: https://cloud.google.com/dataproc/docs/resources/faq#what_type_of_jobs_can_i_run


NEW QUESTION # 168
To run a TensorFlow training job on your own computer using Cloud Machine Learning Engine, what would your command start with?

  • A. gcloud ml-engine jobs submit training local
  • B. gcloud ml-engine local train
  • C. You can't run a TensorFlow program on your own computer using Cloud ML Engine .
  • D. gcloud ml-engine jobs submit training

Answer: B

Explanation:
Explanation
gcloud ml-engine local train - run a Cloud ML Engine training job locally This command runs the specified module in an environment similar to that of a live Cloud ML Engine Training Job.
This is especially useful in the case of testing distributed models, as it allows you to validate that you are properly interacting with the Cloud ML Engine cluster configuration.
Reference: https://cloud.google.com/sdk/gcloud/reference/ml-engine/local/train


NEW QUESTION # 169
You need to look at BigQuery data from a specific table multiple times a day. The underlying table you are querying is several petabytes in size, but you want to filter your data and provide simple aggregations to downstream users. You want to run queries faster and get up-to-date insights quicker. What should you do?

  • A. Use a cached query to accelerate time to results.
  • B. Run a scheduled query to pull the necessary data at specific intervals daily.
  • C. Create a materialized view based off of the query being run.
  • D. Limit the query columns being pulled in the final result.

Answer: C

Explanation:
Materialized views are precomputed views that periodically cache the results of a query for increased performance and efficiency. BigQuery leverages precomputed results from materialized views and whenever possible reads only changes from the base tables to compute up-to-date results. Materialized views can significantly improve the performance of workloads that have the characteristic of common and repeated queries. Materialized views can also optimize queries with high computation cost and small dataset results, such as filtering and aggregating large tables. Materialized views are refreshed automatically when the base tables change, so they always return fresh data. Materialized views can also be used by the BigQuery optimizer to process queries to the base tables, if any part of the query can be resolved by querying the materialized view. References:
* Introduction to materialized views
* Create materialized views
* BigQuery Materialized View Simplified: Steps to Create and 3 Best Practices
* Materialized view in Bigquery


NEW QUESTION # 170
You have a BigQuery table that ingests data directly from a Pub/Sub subscription. The ingested data is encrypted with a Google-managed encryption key. You need to meet a new organization policy that requires you to use keys from a centralized Cloud Key Management Service (Cloud KMS) project to encrypt data at rest. What should you do?

  • A. Create a new BigOuory table by using customer-managed encryption keys (CMEK), and migrate the data from the old BigQuery table.
  • B. Use Cloud KMS encryption key with Dataflow to ingest the existing Pub/Sub subscription to the existing BigQuery table.
  • C. Create a new BigOuery table and Pub/Sub topic by using customer-managed encryption keys (CMEK), and migrate the data from the old Bigauery table.
  • D. Create a new Pub/Sub topic with CMEK and use the existing BigQuery table by using Google-managed encryption key.

Answer: A

Explanation:
To use CMEK for BigQuery, you need to create a key ring and a key in Cloud KMS, and then specify the key resource name when creating or updating a BigQuery table. You cannot change the encryption type of an existing table, so you need to create a new table with CMEK and copy the data from the old table with Google-managed encryption key.
References:
* Customer-managed Cloud KMS keys | BigQuery | Google Cloud
* Creating and managing encryption keys | Cloud KMS Documentation | Google Cloud


NEW QUESTION # 171
Which of the following is NOT true about Dataflow pipelines?

  • A. Dataflow pipelines can be programmed in Java
  • B. Dataflow pipelines are tied to Dataflow, and cannot be run on any other runner
  • C. Dataflow pipelines can consume data from other Google Cloud services
  • D. Dataflow pipelines use a unified programming model, so can work both with streaming and batch data sources

Answer: B

Explanation:
Explanation
Dataflow pipelines can also run on alternate runtimes like Spark and Flink, as they are built using the Apache Beam SDKs Reference: https://cloud.google.com/dataflow/


NEW QUESTION # 172
You are working on a sensitive project involving private user dat
a. You have set up a project on Google Cloud Platform to house your work internally. An external consultant is going to assist with coding a complex transformation in a Google Cloud Dataflow pipeline for your project. How should you maintain users' privacy?

  • A. Grant the consultant the Cloud Dataflow Developer role on the project.
  • B. Create a service account and allow the consultant to log on with it.
  • C. Create an anonymized sample of the data for the consultant to work with in a different project.
  • D. Grant the consultant the Viewer role on the project.

Answer: B


NEW QUESTION # 173
You need to choose a database to store time series CPU and memory usage for millions of computers. You need to store this data in one-second interval samples. Analysts will be performing real-time, ad hoc analytics against the database. You want to avoid being charged for every query executed and ensure that the schema design will allow for future growth of the dataset. Which database and data model should you choose?

  • A. Create a wide table in Cloud Bigtable with a row key that combines the computer identifier with the sample time at each minute, and combine the values for each second as column data.
  • B. Create a table in BigQuery, and append the new samples for CPU and memory to the table
  • C. Create a narrow table in Cloud Bigtable with a row key that combines the Computer Engine computer identifier with the sample time at each second
  • D. Create a wide table in BigQuery, create a column for the sample value at each second, and update the row with the interval for each second

Answer: A


NEW QUESTION # 174
The YARN ResourceManager and the HDFS NameNode interfaces are available on a Cloud Dataproc cluster ____.

  • A. worker node
  • B. application node
  • C. master node
  • D. conditional node

Answer: C

Explanation:
The YARN ResourceManager and the HDFS NameNode interfaces are available on a Cloud Dataproc cluster master node. The cluster master-host-name is the name of your Cloud Dataproc cluster followed by an -m suffix-for example, if your cluster is named "my-cluster", the master-host-name would be "my-cluster-m".


NEW QUESTION # 175
Case Study: 2 - MJTelco
Company Overview
MJTelco is a startup that plans to build networks in rapidly growing, underserved markets around the world. The company has patents for innovative optical communications hardware. Based on these patents, they can create many reliable, high-speed backbone links with inexpensive hardware.
Company Background
Founded by experienced telecom executives, MJTelco uses technologies originally developed to overcome communications challenges in space. Fundamental to their operation, they need to create a distributed data infrastructure that drives real-time analysis and incorporates machine learning to continuously optimize their topologies. Because their hardware is inexpensive, they plan to overdeploy the network allowing them to account for the impact of dynamic regional politics on location availability and cost. Their management and operations teams are situated all around the globe creating many-to- many relationship between data consumers and provides in their system. After careful consideration, they decided public cloud is the perfect environment to support their needs.
Solution Concept
MJTelco is running a successful proof-of-concept (PoC) project in its labs. They have two primary needs:
Scale and harden their PoC to support significantly more data flows generated when they ramp to more than 50,000 installations.
Refine their machine-learning cycles to verify and improve the dynamic models they use to control topology definition.
MJTelco will also use three separate operating environments ?development/test, staging, and production ?
to meet the needs of running experiments, deploying new features, and serving production customers.
Business Requirements
Scale up their production environment with minimal cost, instantiating resources when and where needed in an unpredictable, distributed telecom user community. Ensure security of their proprietary data to protect their leading-edge machine learning and analysis.
Provide reliable and timely access to data for analysis from distributed research workers Maintain isolated environments that support rapid iteration of their machine-learning models without affecting their customers.
Technical Requirements
Ensure secure and efficient transport and storage of telemetry data Rapidly scale instances to support between 10,000 and 100,000 data providers with multiple flows each.
Allow analysis and presentation against data tables tracking up to 2 years of data storing approximately
100m records/day
Support rapid iteration of monitoring infrastructure focused on awareness of data pipeline problems both in telemetry flows and in production learning cycles.
CEO Statement
Our business model relies on our patents, analytics and dynamic machine learning. Our inexpensive hardware is organized to be highly reliable, which gives us cost advantages. We need to quickly stabilize our large distributed data pipelines to meet our reliability and capacity commitments.
CTO Statement
Our public cloud services must operate as advertised. We need resources that scale and keep our data secure. We also need environments in which our data scientists can carefully study and quickly adapt our models. Because we rely on automation to process our data, we also need our development and test environments to work as we iterate.
CFO Statement
The project is too large for us to maintain the hardware and software required for the data and analysis.
Also, we cannot afford to staff an operations team to monitor so many data feeds, so we will rely on automation and infrastructure. Google Cloud's machine learning will allow our quantitative researchers to work on our high-value problems instead of problems with our data pipelines.
You create a new report for your large team in Google Data Studio 360. The report uses Google BigQuery as its data source. It is company policy to ensure employees can view only the data associated with their region, so you create and populate a table for each region. You need to enforce the regional access policy to the data.
Which two actions should you take? (Choose two.)

  • A. Ensure each table is included in a dataset for a region.
  • B. Adjust the settings for each view to allow a related region-based security group view access.
  • C. Adjust the settings for each table to allow a related region-based security group view access.
  • D. Adjust the settings for each dataset to allow a related region-based security group view access.
  • E. Ensure all the tables are included in global dataset.

Answer: A,B


NEW QUESTION # 176
......

Most Reliable Google Professional-Data-Engineer Training Materials: https://actual4test.practicetorrent.com/Professional-Data-Engineer-practice-exam-torrent.html