Examers.io
ExamsOrganizationsHow it worksPricingHelp & FAQ
CompTIA logo

CompTIADataX

Domain 2Objective 4

Model Iteration DY0-001 Practice Questions (Page 1)

Part of the Modeling, analysis, and outcomes domain, which accounts for 24% of the DY0-001 exam. CompTIA does not publish an official question count, but from its 165-minute exam (~65–110 total, ~16–26 in this domain), expect 3–5 from this objective — we provide 13 practice questions to prepare you well beyond it. (estimate)

13questions here
3free pages
4concepts
24%of the exam

Questions 1–5

  1. 1foundation · easy

    A data scientist has trained a model and wants to ensure it generalizes well to unseen data. Which validation method is most appropriate to estimate the model's performance on new data?

    Select an answer first
  2. 2foundation · easy

    When designing a data model for a fraud detection system, the data scientist must ensure that the model can identify rare fraudulent transactions without overwhelming the team with false alarms. Which design principle is most important to address?

    Select an answer first
  3. 3application · medium

    A data science team is building a model to predict customer churn. They have a dataset with 50,000 records and 30 features. The team has split the data into training (60%), validation (20%), and test (20%) sets. After training a model, they evaluate it on the validation set and achieve an AUC of 0.82. They then use the validation set to tune hyperparameters and achieve an AUC of 0.85. Finally, they evaluate the model on the test set and achieve an AUC of 0.80. What is the most likely explanation for the drop in performance from validation to test?

    Select an answer first
  4. 4application · medium

    A data science team is developing a fraud detection model. They split the data into training (70%), validation (15%), and test (15%) sets. After several iterations, the model achieves 99.8% accuracy on the training set and 99.5% on the validation set. However, on the test set, the accuracy drops to 97.2%. The team is considering deploying the model. What should the team do first?

    Select an answer first
  5. 5foundation · easy

    A data scientist must choose between a linear regression model and a random forest model for a dataset with 10,000 samples and 50 features. The primary objective is to maximize predictive accuracy, and the dataset contains both linear and non-linear relationships. Which criterion is most important for selecting the model?

    Select an answer first
Finished these 5 questions?

Review the revealed explanations, or continue through the curriculum.

Free Basic Practice is a study aid with revealable answers — not a scored exam. Examers.io is independent and not affiliated with or endorsed by CompTIA. “DY0-001” is a trademark of its owner, used for identification only.