Databricks-Machine-Learning-Associate: Databricks Certified Machine Learning Associate Scientist Practice Questions
The free Databricks-Machine-Learning-Associate: Databricks Certified Machine Learning Associate questions that deal with scientist, with answers and explanations. The full bank and the timed practice test cover every topic the exam asks about.
Question #4
A data scientist is using Spark ML to engineer features for an exploratory machine learning project. They decide they want to standardize their features using the following code block: Upon code review, a colleague expressed concern with the features being standardized prior to splitting the data into a training set and a test set. Which of the following changes can the data scientist make to address the concern?

Correct answer: E
Explanation
To address the concern about standardizing features prior to splitting the data, the correct approach is to use the Pipeline API to ensure that only the training data's summary statistics are used to standardize the test data. This is achieved by fitting the StandardScaler (or any scaler) on the training data and then transforming both the training and test data using the fitted scaler. This approach prevents information leakage from the test data into the model training process and ensures that the model is evaluated fairly. References: • Best Practices in Preprocessing in Spark ML (Handling Data Splits and Feature Standardization).
Question #8
A data scientist has created a linear regression model that useslog(price)as a label variable. Using this model, they have performed inference and the predictions and actual label values are in Spark DataFramepreds_df. They are using the following code block to evaluate the model: regression_evaluator.setMetricName("rmse").evaluate(preds_df) Which of the following changes should the data scientist make to evaluate the RMSE in a way that is comparable withprice?
Correct answer: D
Explanation
When evaluating the RMSE for a model that predicts log-transformed prices, the predictions need to be transformed back to the original scale to obtain an RMSE that is comparable with the actual price values. This is done by exponentiating the predictions before computing the RMSE. The RMSE should be computed on the same scale as the original data to provide a meaningful measure of error. References: • Databricks documentation on regression evaluation: Regression Evaluation
Continue with Databricks-Machine-Learning-Associate: Databricks Certified Machine Learning Associate Exam
Unlock the full question bank
You have read the first 10 questions. A subscription opens every question in Databricks-Machine-Learning-Associate: Databricks Certified Machine Learning Associate Exam, the full timed practice test, and your progress and weak-topic reporting.
Single exam
$19.99for 30 days
Full question bank and practice test for one exam, for 30 days.
Single exam
$49.99for 1 year
One exam for a full year. Nothing renews and nothing to cancel.
Full access
$39.99/mo
Every exam in the catalogue, month to month.
Full access
$199.99/yr
Every exam in the catalogue for a year.
Already subscribed? Sign in to pick up where you left off.
