Examers.io
ExamsOrganizationsHow it worksPricingHelp & FAQ
DATABRICKS

Databricks Certified Data Engineer Professional

DATA-ENGINEER-PROFESSIONALData Engineer Professional

The Databricks Certified Data Engineer Professional certification validates your advanced skills in building, optimizing, and maintaining production-grade data engineering solutions on the Databricks Data Intelligence Platform. It is designed for experienced data engineers who architect secure, reliable, and cost-effective ETL pipelines, implement streaming workloads, and orchestrate workflows using Python and SQL. Earning this credential demonstrates your ability to deliver complex data engineering projects at scale.

639 practice questions · Updated 2026-07-30

10Domains
26Objectives
199Concepts
639Questions

DATA-ENGINEER-PROFESSIONAL Curriculum

Every domain, objective, and concept the DATA-ENGINEER-PROFESSIONAL exam measures.

Using Python and Tools for development

7 concepts · 23 questions
  1. Scalable Python Project Structure for Bundles
  2. Bundle Configuration and Deployment Automation
  3. Managing PyPI Packages in Databricks
  4. Installing Local Wheels and Source Archives
  5. Troubleshooting Dependency Conflicts
  6. Pandas UDF Fundamentals
  7. Python UDF Development
  1. Lakeflow Spark Declarative Pipelines fundamentals
  2. Autoloader for incremental data ingestion
  3. Managing pipeline lifecycle
  4. Automating ETL with Jobs
  5. Streaming tables vs materialized views
  6. AUTO CDC APIs for change data capture
  7. Choosing between Structured Streaming and Lakeflow Pipelines
  8. Control flow operators in pipelines
  9. Pipeline configuration for environments and dependencies
  10. Unit and integration testing for pipelines
  11. Using the built-in debugger for pipeline code

  1. Data Format Characteristics
  2. Format Selection Criteria
  3. Delta Lake Ingestion
  4. Parquet and ORC Ingestion
  5. Avro Ingestion
  6. JSON Ingestion
  7. CSV Ingestion
  8. XML Ingestion
  9. Text Ingestion
  10. Binary Ingestion
  11. Message Bus Ingestion
  12. Cloud Storage Ingestion
  13. Incremental Ingestion Patterns
  14. Schema and Data Validation
  15. Performance Optimization for Ingestion
  1. Delta append-only pipelines
  2. Batch and streaming ingestion
  3. Structured Streaming with Delta
  4. Schema enforcement and evolution
  5. Idempotent writes
  6. Delta table properties for append-only
  7. Handling late and out-of-order data
  8. Monitoring and checkpointing

  1. Window Functions
  2. Advanced Joins
  3. Aggregations
  4. Optimizing Transformations
  5. PySpark DataFrame API
  6. Spark SQL Syntax
  1. Quarantine bad data in Lakeflow Spark Declarative Pipelines
  2. Quarantine bad data with Autoloader in classic jobs
  3. Define bad data criteria for quarantine
  4. Implement quarantine workflow with error handling
  5. Monitor and manage quarantined data

  1. Delta Sharing Architecture
  2. Databricks-to-Databricks Sharing (D2D)
  3. Open Sharing Protocol (D2O)
  4. Security and Access Control
  5. Sharing Data Assets
  6. Recipient Management
  7. Troubleshooting and Monitoring
  1. Lakehouse Federation Overview
  2. Supported Source Systems
  3. Creating Foreign Catalogs
  4. Managing Foreign Connections
  5. Querying Federated Data
  6. Governance with Unity Catalog
  7. Security and Credential Management
  8. Performance Considerations
  9. Troubleshooting and Monitoring
  1. Delta Sharing Overview
  2. Delta Sharing Architecture
  3. Creating a Delta Share
  4. Managing Recipients
  5. Sharing Live Data
  6. Accessing Shared Data
  7. Security and Permissions

Monitoring

8 concepts · 27 questions
  1. System Tables Overview
  2. Querying System Tables
  3. Interpreting System Table Data
  4. Query Profiler UI Usage
  5. Spark UI Usage
  6. Databricks REST API for Monitoring
  7. Databricks CLI for Monitoring
  8. Lakeflow Spark Declarative Pipelines Event Logs

Alerting

6 concepts · 16 questions
  1. SQL Alert Creation
  2. SQL Alert Management
  3. SQL Alert Notifications
  4. Job Status Notifications via UI
  5. Job Performance Notifications via UI
  6. Job Notifications via Jobs API

  1. Managed vs External Tables
  2. Automatic Lifecycle Management
  3. Centralized Governance
  4. No Manual Location Management
  5. Built-in Data Lineage
  6. Simplified Schema Evolution
  7. Reduced Storage Cleanup Effort
  8. Consistent Metadata Handling
  1. Deletion Vectors Overview
  2. Deletion Vectors Mechanics
  3. Deletion Vectors Benefits and Trade-offs
  4. Deletion Vectors Configuration and Usage
  5. Liquid Clustering Overview
  6. Liquid Clustering Mechanics
  7. Liquid Clustering Benefits and Trade-offs
  8. Liquid Clustering Configuration and Usage
  9. Comparison and Selection
  1. Data Skipping
  2. File Pruning
  3. Z-Ordering
  4. Partitioning vs. Z-Ordering
  5. Statistics Collection
  6. Optimized Writes
  7. Auto Optimize
  8. Bloom Filter Indexing
  1. CDF limitations in streaming tables
  2. CDF for latency enhancement
  3. CDF configuration in Delta Lake
  4. CDF read options for streaming
  5. CDF vs. streaming table trade-offs
  1. Query Profile Overview
  2. Identifying Bottlenecks
  3. Data Skipping Analysis
  4. Join Type Efficiency
  5. Data Shuffling Detection

Applying Data Security mechanisms.

8 concepts · 23 questions
  1. ACL Basics
  2. Workspace Object Permissions
  3. Principle of Least Privilege
  4. Policy Enforcement with ACLs
  5. Row Filters
  6. Column Masks
  7. Anonymization Techniques
  8. Pseudonymization Techniques

Ensuring Compliance

8 concepts · 28 questions
  1. PII Detection in Pipelines
  2. Dynamic Masking Techniques
  3. Streaming Masking Implementation
  4. Batch Masking Implementation
  5. Compliance-Driven Pipeline Design
  6. Data Retention Policy Interpretation
  7. Data Purging Implementation
  8. Purge Verification and Auditing

  1. Understanding Databricks Unity Catalog metadata
  2. Adding table and column comments
  3. Setting table properties
  4. Managing schema descriptions
  5. Using the Catalog Explorer for metadata
  6. Querying metadata with DESCRIBE and SHOW
  7. Documenting data lineage and ownership
  1. Unity Catalog permission inheritance hierarchy
  2. Inheritance of privileges from parent to child objects
  3. Overriding inherited permissions
  4. Effect of deny rules on inheritance
  5. Permission inheritance for tables and views
  6. Permission inheritance for functions and models
  7. Permission inheritance for volumes and external locations
  8. Inheritance of ownership
  9. Best practices for managing permission inheritance

Debugging and Troubleshooting

9 concepts · 34 questions
  1. Spark UI navigation for diagnostics
  2. Cluster log analysis
  3. System table queries for job metadata
  4. Query profile interpretation
  5. Error analysis and remediation
  6. Job repairs for failed runs
  7. Parameter overrides for job runs
  8. Lakeflow Spark Declarative Pipelines event logs
  9. Spark UI for pipeline debugging

Deploying CI/CD

7 concepts · 25 questions
  1. Declarative Automation Bundles basics
  2. Bundle configuration files
  3. Deploying resources with bundles
  4. Bundle lifecycle and validation
  5. Git-based CI/CD integration
  6. CI/CD workflow configuration
  7. Automated deployment in CI/CD

  1. Delta Lake table design for large datasets
  2. Delta Lake schema evolution and enforcement
  3. Delta Lake time travel and versioning
  4. Delta Lake optimization techniques
  5. Delta Lake data skipping and statistics
  6. Delta Lake transaction and concurrency control
  7. Delta Lake table maintenance and lifecycle
  1. Liquid Clustering Overview
  2. Clustering Keys
  3. Automatic Layout Management
  4. Query Performance Optimization
  5. Table Creation with Liquid Clustering
  6. Altering Clustering Keys
  7. Compatibility and Limitations
  1. Liquid Clustering Fundamentals
  2. Partitioning Limitations
  3. ZOrder Limitations
  4. Benefits of Liquid Clustering
  5. Comparison with Partitioning and ZOrder
  6. Use Cases for Liquid Clustering
  1. Dimensional Modeling Fundamentals
  2. Star Schema Design
  3. Snowflake Schema Design
  4. Fact Table Types
  5. Dimension Table Types
  6. Granularity and Aggregation
  7. Query Optimization in Dimensional Models
Ready to practice?Test your knowledge with exam-style questions or take an intelligent quiz tailored to your level.

Percentages reflect share of the current practice bank, not official exam weightings — no structured per-skill weight is published for DATA-ENGINEER-PROFESSIONAL, so none is invented.