
Databricks Certified Data Engineer Professional
The Databricks Certified Data Engineer Professional certification validates your advanced skills in building, optimizing, and maintaining production-grade data engineering solutions on the Databricks Data Intelligence Platform. It is designed for experienced data engineers who architect secure, reliable, and cost-effective ETL pipelines, implement streaming workloads, and orchestrate workflows using Python and SQL. Earning this credential demonstrates your ability to deliver complex data engineering projects at scale.
639 practice questions · Updated 2026-07-30
DATA-ENGINEER-PROFESSIONAL Curriculum
Every domain, objective, and concept the DATA-ENGINEER-PROFESSIONAL exam measures.
- Scalable Python Project Structure for Bundles
- Bundle Configuration and Deployment Automation
- Managing PyPI Packages in Databricks
- Installing Local Wheels and Source Archives
- Troubleshooting Dependency Conflicts
- Pandas UDF Fundamentals
- Python UDF Development
- Lakeflow Spark Declarative Pipelines fundamentals
- Autoloader for incremental data ingestion
- Managing pipeline lifecycle
- Automating ETL with Jobs
- Streaming tables vs materialized views
- AUTO CDC APIs for change data capture
- Choosing between Structured Streaming and Lakeflow Pipelines
- Control flow operators in pipelines
- Pipeline configuration for environments and dependencies
- Unit and integration testing for pipelines
- Using the built-in debugger for pipeline code
- Data Format Characteristics
- Format Selection Criteria
- Delta Lake Ingestion
- Parquet and ORC Ingestion
- Avro Ingestion
- JSON Ingestion
- CSV Ingestion
- XML Ingestion
- Text Ingestion
- Binary Ingestion
- Message Bus Ingestion
- Cloud Storage Ingestion
- Incremental Ingestion Patterns
- Schema and Data Validation
- Performance Optimization for Ingestion
- Delta append-only pipelines
- Batch and streaming ingestion
- Structured Streaming with Delta
- Schema enforcement and evolution
- Idempotent writes
- Delta table properties for append-only
- Handling late and out-of-order data
- Monitoring and checkpointing
- Window Functions
- Advanced Joins
- Aggregations
- Optimizing Transformations
- PySpark DataFrame API
- Spark SQL Syntax
- Quarantine bad data in Lakeflow Spark Declarative Pipelines
- Quarantine bad data with Autoloader in classic jobs
- Define bad data criteria for quarantine
- Implement quarantine workflow with error handling
- Monitor and manage quarantined data
- Delta Sharing Architecture
- Databricks-to-Databricks Sharing (D2D)
- Open Sharing Protocol (D2O)
- Security and Access Control
- Sharing Data Assets
- Recipient Management
- Troubleshooting and Monitoring
- Lakehouse Federation Overview
- Supported Source Systems
- Creating Foreign Catalogs
- Managing Foreign Connections
- Querying Federated Data
- Governance with Unity Catalog
- Security and Credential Management
- Performance Considerations
- Troubleshooting and Monitoring
- System Tables Overview
- Querying System Tables
- Interpreting System Table Data
- Query Profiler UI Usage
- Spark UI Usage
- Databricks REST API for Monitoring
- Databricks CLI for Monitoring
- Lakeflow Spark Declarative Pipelines Event Logs
- SQL Alert Creation
- SQL Alert Management
- SQL Alert Notifications
- Job Status Notifications via UI
- Job Performance Notifications via UI
- Job Notifications via Jobs API
- Managed vs External Tables
- Automatic Lifecycle Management
- Centralized Governance
- No Manual Location Management
- Built-in Data Lineage
- Simplified Schema Evolution
- Reduced Storage Cleanup Effort
- Consistent Metadata Handling
- Deletion Vectors Overview
- Deletion Vectors Mechanics
- Deletion Vectors Benefits and Trade-offs
- Deletion Vectors Configuration and Usage
- Liquid Clustering Overview
- Liquid Clustering Mechanics
- Liquid Clustering Benefits and Trade-offs
- Liquid Clustering Configuration and Usage
- Comparison and Selection
- Data Skipping
- File Pruning
- Z-Ordering
- Partitioning vs. Z-Ordering
- Statistics Collection
- Optimized Writes
- Auto Optimize
- Bloom Filter Indexing
- CDF limitations in streaming tables
- CDF for latency enhancement
- CDF configuration in Delta Lake
- CDF read options for streaming
- CDF vs. streaming table trade-offs
- Query Profile Overview
- Identifying Bottlenecks
- Data Skipping Analysis
- Join Type Efficiency
- Data Shuffling Detection
- ACL Basics
- Workspace Object Permissions
- Principle of Least Privilege
- Policy Enforcement with ACLs
- Row Filters
- Column Masks
- Anonymization Techniques
- Pseudonymization Techniques
- PII Detection in Pipelines
- Dynamic Masking Techniques
- Streaming Masking Implementation
- Batch Masking Implementation
- Compliance-Driven Pipeline Design
- Data Retention Policy Interpretation
- Data Purging Implementation
- Purge Verification and Auditing
- Understanding Databricks Unity Catalog metadata
- Adding table and column comments
- Setting table properties
- Managing schema descriptions
- Using the Catalog Explorer for metadata
- Querying metadata with DESCRIBE and SHOW
- Documenting data lineage and ownership
- Unity Catalog permission inheritance hierarchy
- Inheritance of privileges from parent to child objects
- Overriding inherited permissions
- Effect of deny rules on inheritance
- Permission inheritance for tables and views
- Permission inheritance for functions and models
- Permission inheritance for volumes and external locations
- Inheritance of ownership
- Best practices for managing permission inheritance
- Spark UI navigation for diagnostics
- Cluster log analysis
- System table queries for job metadata
- Query profile interpretation
- Error analysis and remediation
- Job repairs for failed runs
- Parameter overrides for job runs
- Lakeflow Spark Declarative Pipelines event logs
- Spark UI for pipeline debugging
- Declarative Automation Bundles basics
- Bundle configuration files
- Deploying resources with bundles
- Bundle lifecycle and validation
- Git-based CI/CD integration
- CI/CD workflow configuration
- Automated deployment in CI/CD
- Delta Lake table design for large datasets
- Delta Lake schema evolution and enforcement
- Delta Lake time travel and versioning
- Delta Lake optimization techniques
- Delta Lake data skipping and statistics
- Delta Lake transaction and concurrency control
- Delta Lake table maintenance and lifecycle
- Liquid Clustering Overview
- Clustering Keys
- Automatic Layout Management
- Query Performance Optimization
- Table Creation with Liquid Clustering
- Altering Clustering Keys
- Compatibility and Limitations
- Liquid Clustering Fundamentals
- Partitioning Limitations
- ZOrder Limitations
- Benefits of Liquid Clustering
- Comparison with Partitioning and ZOrder
- Use Cases for Liquid Clustering
- Dimensional Modeling Fundamentals
- Star Schema Design
- Snowflake Schema Design
- Fact Table Types
- Dimension Table Types
- Granularity and Aggregation
- Query Optimization in Dimensional Models
Percentages reflect share of the current practice bank, not official exam weightings — no structured per-skill weight is published for DATA-ENGINEER-PROFESSIONAL, so none is invented.