
DatabricksCertified Machine Learning Associate
Domain 2Objective 5
Compare and Contrast Imputing Missing Values with the Mean or Median or Mode Value MACHINE-LEARNING-ASSOCIATE Practice Questions (Page 2)
Part of the Section 2: Data Processing domain, which makes up ~27% of our current practice bank. Databricks does not publish an official question count, but from its 90-minute exam (~35–60 total, ~9–16 in this domain), expect 1–2 from this objective — we provide 14 practice questions to prepare you well beyond it. (estimate)
14questions here
3free pages
5concepts
Questions 6–10
- 6
For which type of data distribution is median imputation particularly recommended over mean imputation?
Select an answer first - 7
Why might mode imputation be a poor choice for a continuous numerical feature in a machine learning model?
Select an answer first - 8
A machine learning engineer is working on a churn prediction model. The dataset includes a categorical feature 'customer_segment' with values 'Basic', 'Standard', and 'Premium'. About 5% of the records have this field missing. The class distribution is: Basic 70%, Standard 20%, Premium 10%. What is the most appropriate imputation strategy for this feature?
Select an answer first - 9
A data scientist is working with a dataset where a numeric feature is normally distributed with no outliers. The missing rate is 5%. They are considering mean imputation. What is the primary advantage of mean imputation in this scenario?
Select an answer first - 10
For which type of missing data is mode imputation most appropriate?
Select an answer first
Finished these 5 questions?
Review the revealed explanations, or continue through the curriculum.
Free Basic Practice is a study aid with revealable answers — not a scored exam. Examers.io is independent and not affiliated with or endorsed by Databricks. “MACHINE-LEARNING-ASSOCIATE” is a trademark of its owner, used for identification only.