
DatabricksCertified Machine Learning Associate
Domain 3Objective 3
Training Pipelines and Cross-Validation MACHINE-LEARNING-ASSOCIATE Practice Questions (Page 4)
Part of the Section 3: Model Development domain, which makes up ~23% of our current practice bank. Databricks does not publish an official question count, but from its 90-minute exam (~35–60 total, ~8–14 in this domain), expect 2–4 from this objective — we provide 29 practice questions to prepare you well beyond it. (estimate)
29questions here
6free pages
9concepts
Questions 16–20
- 16
Which of the following correctly lists the typical stages of a training pipeline in the order they are executed?
Select an answer first - 17
A team is working with a highly imbalanced classification dataset where the positive class is only 2% of the data. They want to use cross-validation to evaluate a model. What is the most appropriate cross-validation strategy?
Select an answer first - 18
A team is comparing two classification models. They have a small dataset of 2,000 rows and want a performance estimate that is less sensitive to which particular rows end up in the validation set. They also want to see per-fold scores to understand variability. Which approach should they use?
Select an answer first - 19
A data scientist wants to evaluate a logistic regression model using cross-validation. They also want to capture multiple metrics (accuracy, precision, recall) and the training time for each fold. Which scikit-learn function should they use?
Select an answer first - 20
Which of the following is a complexity introduced by k-fold cross-validation that is not present in a single train-validation split?
Select an answer first
Finished these 5 questions?
Review the revealed explanations, or continue through the curriculum.
Free Basic Practice is a study aid with revealable answers — not a scored exam. Examers.io is independent and not affiliated with or endorsed by Databricks. “MACHINE-LEARNING-ASSOCIATE” is a trademark of its owner, used for identification only.