Examers.io
ExamsOrganizationsHow it worksPricingHelp & FAQ
Databricks logo

DatabricksCertified Machine Learning Associate

Domain 2Objective 3

Create Visualizations for Categorical or Continuous Features MACHINE-LEARNING-ASSOCIATE Practice Questions (Page 2)

Part of the Section 2: Data Processing domain, which makes up ~27% of our current practice bank. Databricks does not publish an official question count, but from its 90-minute exam (~35–60 total, ~9–16 in this domain), expect 1–2 from this objective — we provide 11 practice questions to prepare you well beyond it. (estimate)

11questions here
3free pages
2concepts

Questions 6–10

  1. 6application · medium

    A data analyst is exploring a dataset of patient ages. The `age` column is continuous. The analyst also has a `gender` column with values 'Male' and 'Female'. The analyst wants to compare the distribution of ages between the two gender groups. Which visualization is the most appropriate?

    Select an answer first
  2. 7application · medium

    A machine learning engineer is working with a dataset of housing prices. The `price` column is a continuous variable with a heavily right-skewed distribution. The `house_type` column is categorical with values 'Apartment', 'House', and 'Townhouse'. The engineer wants to visualize the distribution of `price` for each `house_type` to check for outliers and compare their spreads. Which visualization is the most effective for this task?

    Select an answer first
  3. 8application · easy

    A data scientist is building a model to predict customer lifetime value. The dataset includes a categorical feature `subscription_type` with values 'Basic', 'Premium', and 'Enterprise'. The scientist wants to check the class balance of this feature to understand if there is a significant class imbalance. Which visualization is most appropriate?

    Select an answer first
  4. 9application · easy

    A data analyst is preparing a report on website traffic sources. The `source` column contains values like 'Organic Search', 'Paid Ads', 'Social Media', and 'Direct'. The analyst wants to show the proportion of total sessions attributed to each source. Which visualization is the most appropriate?

    Select an answer first
  5. 10expert · hard

    A data scientist is analyzing a dataset with a continuous feature `income` that is heavily right-skewed. The scientist wants to compare the distribution of `income` across two categorical groups, `urban` and `rural`. The scientist is concerned that the skewness will make it difficult to see the differences in the bulk of the data. Which approach is the most appropriate to address this concern?

    Select an answer first
Finished these 5 questions?

Review the revealed explanations, or continue through the curriculum.

Free Basic Practice is a study aid with revealable answers — not a scored exam. Examers.io is independent and not affiliated with or endorsed by Databricks. “MACHINE-LEARNING-ASSOCIATE” is a trademark of its owner, used for identification only.