Examers.io
ExamsOrganizationsHow it worksPricingHelp & FAQ
DATABRICKS

Databricks Certified Machine Learning Associate

MACHINE-LEARNING-ASSOCIATE

The Databricks Certified Machine Learning Associate certification validates your ability to use Databricks to perform basic machine learning tasks, from data exploration and feature engineering to model training, tuning, evaluation, and deployment. It is designed for machine learning engineers and data scientists who work with Databricks Machine Learning, AutoML, Unity Catalog, and MLflow. Earning this credential demonstrates that you can build and deploy ML models on the Databricks Data Intelligence Platform.

411 practice questions · Updated 2026-07-30

4Domains
23Objectives
124Concepts
411Questions

MACHINE-LEARNING-ASSOCIATE Curriculum

Every domain, objective, and concept the MACHINE-LEARNING-ASSOCIATE exam measures.

MLflow and Model Lifecycle Management

8 concepts · 19 questions
  1. MLflow Client API run search
  2. Manual logging in MLflow
  3. MLflow UI information
  4. Model registration via Unity Catalog
  5. Unity Catalog vs workspace registry benefits
  6. Code promotion vs model promotion
  7. Model tag management
  8. Champion-challenger model promotion

Feature Store Management

6 concepts · 18 questions
  1. Account-level vs workspace-level feature store tables
  2. Creating feature store tables in Unity Catalog
  3. Writing data to feature store tables
  4. Training models with feature store features
  5. Scoring models with feature store features
  6. Online vs offline feature tables

AutoML and ML Runtime

4 concepts · 21 questions
  1. ML Runtime advantages
  2. AutoML feature selection
  3. AutoML model selection
  4. AutoML development advantages

MLOps Best Practices

2 concepts · 22 questions
  1. MLOps strategy definition
  2. Best practices for MLOps

  1. Spark DataFrame summary() method
  2. dbutils data summaries
  3. Interpreting summary statistics
  1. Standard Deviation Outlier Removal
  2. IQR Outlier Removal
  3. Spark DataFrame Filtering
  4. Window Functions for Percentiles
  1. Categorical feature visualization
  2. Continuous feature visualization
  1. Categorical feature comparison methods
  2. Continuous feature comparison methods
  3. Assumptions and conditions for each method
  4. Interpreting results of feature comparisons
  1. Mean imputation
  2. Median imputation
  3. Mode imputation
  4. Comparison of imputation methods
  5. Impact on model performance
  1. Identify missing values
  2. Understand imputation methods
  3. Apply mode imputation
  4. Apply mean imputation
  5. Apply median imputation
  6. Implement imputation in code
  1. One-hot encoding fundamentals
  2. Implementation in Spark ML
  3. Handling categorical columns
  4. Integration with pipelines
  5. Impact on model performance
  1. One-hot encoding definition
  2. Appropriate model types for one-hot encoding
  3. Inappropriate model types for one-hot encoding
  4. Appropriate datasets for one-hot encoding
  5. Inappropriate datasets for one-hot encoding
  1. Log scale transformation definition
  2. Purpose of log transformation
  3. Identifying right-skewed data
  4. Handling positive data requirement
  5. Effect on model performance
  6. Interpreting results after log transformation

Model Selection and Evaluation

7 concepts · 22 questions
  1. Algorithm Selection Based on ML Foundations
  2. Comparing Estimators and Transformers
  3. Classification Metrics
  4. Regression Metrics
  5. Metric Selection for Scenario Objectives
  6. Exponentiating Log-Transformed Variables
  7. Model Complexity and Bias-Variance Tradeoff

Data Preparation and Imbalance

4 concepts · 21 questions
  1. Data imbalance definition
  2. Resampling techniques
  3. Class weighting
  4. Evaluation metrics for imbalanced data

Training Pipelines and Cross-Validation

9 concepts · 29 questions
  1. Training Pipeline Components
  2. Pipeline Implementation
  3. Train-Validation Split
  4. Cross-Validation Overview
  5. Benefits of Cross-Validation
  6. Downsides of Cross-Validation
  7. Performing Cross-Validation
  8. Grid Search with Cross-Validation
  9. Counting Models in Grid Search with CV

Hyperparameter Tuning

11 concepts · 24 questions
  1. Hyperopt fmin basics
  2. Hyperopt search spaces
  3. Hyperopt objective function
  4. Running fmin optimization
  5. Grid search tuning
  6. Random search tuning
  7. Bayesian search tuning
  8. Hyperopt Bayesian search
  9. Parallelizing single-node models
  10. SparkTrials configuration
  11. Monitoring parallel tuning

  1. Batch Model Serving
  2. Realtime Model Serving
  3. Streaming Model Serving
  4. Advantages of Batch Serving
  5. Advantages of Realtime Serving
  6. Advantages of Streaming Serving
  7. Trade-offs Between Approaches
  8. Selecting an Appropriate Approach
  1. Model endpoint fundamentals
  2. Custom model packaging
  3. Endpoint creation
  4. Endpoint configuration
  5. Endpoint deployment and update
  6. Endpoint invocation
  7. Endpoint monitoring and management

Use pandas to perform batch inference

5 concepts · 15 questions
  1. Pandas DataFrame basics for batch inference
  2. Applying a trained model to a DataFrame
  3. Handling feature preprocessing in batch
  4. Managing model input/output formats
  5. Writing inference results to files
  1. Streaming inference in Delta Live Tables
  2. Defining streaming inference pipelines in DLT
  3. Applying ML models to streaming data in DLT
  4. Handling stateful streaming inference in DLT
  5. Monitoring and managing streaming inference pipelines
  1. Model serving endpoints
  2. Querying endpoints
  3. Endpoint deployment modes
  4. Endpoint management
  1. Data splitting for real-time inference
  2. Endpoint configuration for real-time serving
  3. Handling data skew and distribution
  4. Monitoring and updating data splits
Ready to practice?Test your knowledge with exam-style questions or take an intelligent quiz tailored to your level.

Percentages reflect share of the current practice bank, not official exam weightings — no structured per-skill weight is published for MACHINE-LEARNING-ASSOCIATE, so none is invented.