
Databricks Certified Data Engineer Associate
The Databricks Certified Data Engineer Associate certification validates your ability to use the Databricks Data Intelligence Platform for foundational data engineering tasks. It covers data ingestion, ETL with PySpark, Lakeflow Jobs, CI/CD, and governance. Ideal for data engineers building and maintaining batch and streaming pipelines, this credential demonstrates you can turn raw data into reliable, governed analytics-ready datasets.
623 practice questions · Updated 2026-07-30
7Domains
33Objectives
207Concepts
623Questions
DATA-ENGINEER-ASSOCIATE Curriculum
Every domain, objective, and concept the DATA-ENGINEER-ASSOCIATE exam measures.
- Databricks Architecture Overview
- Core Components of Databricks
- Delta Lake Fundamentals
- Unity Catalog Fundamentals
- Integration of Components
- Databricks compute services overview
- Characteristics of compute services
- Limitations of compute services
- Cost models of compute services
- Selecting compute for workloads
- Batch ingestion patterns
- Streaming ingestion patterns
- Incremental loading patterns
- Loading data from local files
- Lakeflow Connect standard connectors
- Lakeflow Connect managed connectors
- COPY INTO syntax and parameters
- Incremental loading behavior
- Handling schema evolution
- Error handling and validation
- Integration with Unity Catalog
- File format support
- Partitioning and file organization
- Auto Loader overview
- Batch mode operation
- Directory listing mode
- File notification mode
- Schema inference
- Schema enforcement
- Schema evolution
- Schema evolution modes
- Landing data into Unity Catalog tables
- Checkpointing
- Options and configurations
- Lakeflow Connect overview
- Source system configuration
- Data ingestion pipeline setup
- Incremental and full load strategies
- Schema and data type mapping
- Error handling and retry mechanisms
- Monitoring and alerting
- Unity Catalog integration
- Performance optimization
- Security and credential management
- JDBC/ODBC connection setup in notebooks
- Reading data via JDBC/ODBC
- Writing data to cloud storage
- Writing data to Unity Catalog tables
- REST client usage in notebooks
- Orchestrating ingestion with Lakeflow Jobs
- Handling incremental and full loads
- Error handling and retries in ingestion
- Auto Loader capabilities
- Lakeflow Connect standard connectors
- Lakeflow Connect managed connectors
- Partner connectors
- Other ingestion methods
- Selection criteria based on data volume
- Selection criteria based on ingestion frequency
- Selection criteria based on data types
- Governance needs with Unity Catalog
- Identify semi-structured and unstructured data types
- Use Lakeflow Connect for data ingestion
- Use other managed connectors for data ingestion
- Parse and flatten nested JSON data
- Ingest data into Unity Catalog-governed Delta tables
- Handle schema inference and evolution
- Validate and troubleshoot ingestion pipelines
- Reading Bronze Tables with PySpark/SQL
- Handling Null Values in DataFrames
- Standardizing Data Types
- Writing Cleaned Data to Silver Tables
- Inner Join
- Left Join
- Broadcast Join
- Multiple Keys Join
- Cross Join
- Union
- Union All
- Adding columns
- Dropping columns
- Splitting columns
- Renaming columns
- Applying filters
- Exploding arrays
- Manipulating rows
- Manipulating table structures
- DataFrame deduplication
- Aggregation operations
- Count distinct values
- Approximate count distinct
- Summary statistics
- spark.sql.shuffle.partitions
- spark.default.parallelism
- spark.executor.memory and spark.driver.memory
- spark.sql.autoBroadcastJoinThreshold
- Performance measurement and re-measurement
- Gold layer purpose
- Views vs tables
- Materialized views
- Streaming tables
- Building Gold layer objects
- Unity Catalog integration
- BI and analytics optimization
- Data Quality Dimensions
- Validation Rules Definition
- Implementing Checks in Pipelines
- Handling Invalid Records
- Monitoring and Alerting
- Retry policies
- Retry conditions
- Conditional branching
- Looping constructs
- Pipeline orchestration
- Notebook task configuration
- SQL query task configuration
- Dashboard task configuration
- Pipeline task configuration
- Task dependencies in DAG
- DAG-based task graph management
- Scheduled triggers
- File arrival triggers
- Table update triggers
- Trigger type selection
- Time-based triggers
- Data-driven triggers
- Comparing trigger types
- Selecting triggers based on data availability
- Selecting triggers based on pipeline dependencies
- Handling late or missing data
- Best practices for trigger selection
- Accessing Git Folders
- Creating and Switching Branches
- Committing Changes
- Pushing Changes
- Creating Pull Requests
- Managing Git Workflow in UI
- Automation Bundle variables
- Environment-specific overrides
- Promoting code across environments
- Variable resolution precedence
- Bundle configuration files
- Bundle structure and configuration
- Packaging assets
- Environment targeting
- Promotion across environments
- Deployment and validation
- Databricks CLI installation and configuration
- Declarative Automation Bundles overview
- Bundle validation
- Bundle deployment
- Bundle management
- Workspace asset management
- CI/CD integration
- Navigate to the Lakeflow Jobs run history view
- Interpret run history metrics
- Compare current execution times to historical baselines
- Identify trends in job performance
- Interpreting Job Statuses
- Viewing DAG-based Task Graphs
- Tracking Pipeline Run Times
- Tracking Failure Rates
- Monitoring Pipeline Health
- Identify data skew in Spark UI
- Identify excessive shuffling in Spark UI
- Identify disk spilling in Spark UI
- Interpret stage-level metrics in Spark UI
- Liquid Clustering overview
- Clustering keys and data skipping
- Enabling and configuring Liquid Clustering
- Incremental clustering and OPTIMIZE
- Predictive optimization overview
- Enabling and configuring predictive optimization
- Monitoring and benefits
- Cluster startup failure diagnosis
- Library conflict resolution
- Out-of-memory issue diagnosis
- Managed vs external table definitions
- Creating managed and external tables
- Modifying managed and external tables
- Deleting managed and external tables
- Converting between managed and external tables
- Understanding the security hierarchy
- Using the UI to manage privileges
- Using SQL to manage privileges
- Applying GRANT privileges
- Applying REVOKE privileges
- Applying DENY privileges
- Managing privileges for users, groups, and service principals
- Understanding privilege inheritance and precedence
- Column-level masking fundamentals
- Creating masking policies
- Applying masking policies to tables
- Testing and validating masking behavior
- Row-level security fundamentals
- Creating row filters
- Applying row filters to tables
- Testing and validating row-level security
- Combining column masking and row-level security
- Managing policies and filters
- Unity Catalog ABAC overview
- Row-level filtering
- Column masking
- Creating and managing filters and masks
- Applying filters and masks to tables
- Authorization and privileges
- Evaluation and precedence
Ready to practice?Test your knowledge with exam-style questions or take an intelligent quiz tailored to your level.
Percentages reflect share of the current practice bank, not official exam weightings — no structured per-skill weight is published for DATA-ENGINEER-ASSOCIATE, so none is invented.