Databricks-Certified-Professional-Data-Engineer: Databricks Certified Data Engineer Professional Table Practice Questions
The free Databricks-Certified-Professional-Data-Engineer: Databricks Certified Data Engineer Professional questions that deal with table, with answers and explanations. The full bank and the timed practice test cover every topic the exam asks about.
Question #5
A Delta Lake table in the Lakehouse named customer_parsams is used in churn prediction by the machine learning team. The table contains information about customers derived from a number of upstream sources. Currently, the data engineering team populates this table nightly by overwriting the table with the current valid values derived from upstream data sources. Immediately after each update succeeds, the data engineer team would like to determine the difference between the new version and the previous of the table. Given the current implementation, which method can be used?
Correct answer: C
Explanation
Delta Lake provides built-in versioning and time travel capabilities, allowing users to query previous snapshots of a table. This feature is particularly useful for understanding changes between different versions of the table. In this scenario, where the table is overwritten nightly, you can use Delta Lake's time travel feature to execute a query comparing the latest version of the table (the current state) with its previous version. This approach effectively identifies the differences (such as new, updated, or deleted records) between the two versions. The other options do not provide a straightforward or efficient way to directly compare different versions of a Delta Lake table. References: • Delta Lake Documentation on Time Travel: Delta Time Travel • Delta Lake Versioning: Delta Lake Versioning Guide
Question #9
A data engineer is configuring a pipeline that will potentially see late-arriving, duplicate records. In addition to de-duplicating records within the batch, which of the following approaches allows the data engineer to deduplicate data against previously processed records as it is inserted into a Delta table?
Correct answer: C
Explanation
To deduplicate data against previously processed records as it is inserted into a Delta table, you can use the merge operation with an insert-only clause. This allows you to insert new records that do not match any existing records based on a unique key, while ignoring duplicate records that match existing records. For example, you can use the following syntax: MERGE INTO target_table USING source_table ON target_table.unique_key = source_table.unique_key WHEN NOT MATCHED THEN INSERT * This will insert only the records from the source table that have a unique key that is not present in the target table, and skip the records that have a matching key. This way, you can avoid inserting duplicate records into the Delta table. References: • https://docs.databricks.com/delta/delta-update.html#upsert-into-a-table-using- merge • https://docs.databricks.com/delta/delta-update.html#insert-only-merge
Continue with Databricks-Certified-Professional-Data-Engineer: Databricks Certified Data Engineer Professional Exam
Unlock the full question bank
You have read the first 10 questions. A subscription opens every question in Databricks-Certified-Professional-Data-Engineer: Databricks Certified Data Engineer Professional Exam, the full timed practice test, and your progress and weak-topic reporting.
Single exam
$19.99for 30 days
Full question bank and practice test for one exam, for 30 days.
Single exam
$49.99for 1 year
One exam for a full year. Nothing renews and nothing to cancel.
Full access
$39.99/mo
Every exam in the catalogue, month to month.
Full access
$199.99/yr
Every exam in the catalogue for a year.
Already subscribed? Sign in to pick up where you left off.
