
CompTIADataX
Domain 2Objective 4
Model Iteration DY0-001 Practice Questions (Page 2)
Part of the Modeling, analysis, and outcomes domain, which accounts for 24% of the DY0-001 exam. CompTIA does not publish an official question count, but from its 165-minute exam (~65–110 total, ~16–26 in this domain), expect 3–5 from this objective — we provide 13 practice questions to prepare you well beyond it. (estimate)
13questions here
3free pages
4concepts
24%of the exam
Questions 6–10
- 6
A data science team is building a binary classification model to detect spam emails. The dataset has 10,000 emails, with 20% being spam. The team trained a model and is evaluating it on a test set. The confusion matrix shows: True Positives (spam correctly identified) = 150, False Negatives (spam missed) = 50, True Negatives (non-spam correctly identified) = 7,800, False Positives (non-spam flagged as spam) = 2,000. The business wants to minimize the number of legitimate emails that are incorrectly flagged as spam. Which metric should the team focus on?
Select an answer first - 7
A credit risk team is developing a model to predict loan default. They are using a dataset with 100,000 records and 50 features. The team has split the data into training and test sets. They are considering using k-fold cross-validation to evaluate the model. What is the primary benefit of using k-fold cross-validation over a single train-test split?
Select an answer first - 8
A data scientist is designing a model to predict customer churn. The business requires the model to be interpretable so that marketing teams can understand why customers are predicted to leave. Which model design principle is most directly prioritized in this scenario?
Select an answer first - 9
A data science team is developing a model to predict loan default. The dataset has 200,000 records and 100 features. The team has split the data into training (70%) and test (30%) sets. They are using logistic regression and have achieved an AUC of 0.78 on the test set. The business wants to improve the model's performance. The team is considering two approaches: (1) adding polynomial features to capture non-linear relationships, or (2) using a random forest model. The business has a strict requirement that the model must be interpretable for regulatory compliance. Which approach should the team take?
Select an answer first - 10
A data scientist has trained a binary classification model and wants to evaluate its performance on a test set. Which metric is most appropriate when the cost of false negatives is much higher than the cost of false positives?
Select an answer first
Finished these 5 questions?
Review the revealed explanations, or continue through the curriculum.
Free Basic Practice is a study aid with revealable answers — not a scored exam. Examers.io is independent and not affiliated with or endorsed by CompTIA. “DY0-001” is a trademark of its owner, used for identification only.