
DatabricksCertified Machine Learning Professional
Domain 1Objective 2
Scaling and Tuning MACHINE-LEARNING-PROFESSIONAL Practice Questions (Page 4)
Part of the Model Development domain, which accounts for 44% of the MACHINE-LEARNING-PROFESSIONAL exam. Databricks does not publish an official question count, but from its 120-minute exam (~50–80 total, ~22–35 in this domain), expect 6–9 from this objective — we provide 23 practice questions to prepare you well beyond it. (estimate)
23questions here
5free pages
7concepts
44%of the exam
Questions 16–20
- 16
A team is using Optuna to tune hyperparameters for a model. They want to log each trial's parameters and metrics to MLflow. They are running the tuning on a Databricks cluster. What is the recommended way to integrate Optuna with MLflow?
Select an answer first - 17
A team is training a very large transformer model with 100 billion parameters. They have a cluster with 64 GPUs, each with 40 GB memory. The model does not fit on a single GPU. They need to choose between data parallelism and model parallelism. What is the primary consideration?
Select an answer first - 18
A team needs to train a custom reinforcement learning (RL) environment that involves complex simulation and requires frequent communication between the simulation and the model. The training loop is not easily expressible as a Spark job. They are considering using Ray or Spark. Which framework is more suitable?
Select an answer first - 19
Which component of the Spark MLlib pipeline is responsible for transforming raw data into features that can be used by a distributed ML algorithm?
Select an answer first - 20
A company has a Databricks workload that trains a model on a dataset that grows by 10% each month. Currently, the training job runs on a single node with 64 GB RAM and takes 8 hours. The team expects the data to double in the next year. They have a fixed budget and need to keep the training time under 12 hours. They are considering two options: (1) upgrade to a node with 128 GB RAM, or (2) add a second node and use Spark to distribute the training. The model training algorithm is not easily parallelizable. Which option is more appropriate?
Select an answer first
Finished these 5 questions?
Review the revealed explanations, or continue through the curriculum.
Free Basic Practice is a study aid with revealable answers — not a scored exam. Examers.io is independent and not affiliated with or endorsed by Databricks. “MACHINE-LEARNING-PROFESSIONAL” is a trademark of its owner, used for identification only.