Examers.io
ExamsOrganizationsHow it worksPricingHelp & FAQ
Google Cloud logo

Google CloudProfessional Data Engineer

Domain 4Objective 2

4.2 Preparing Data for AI and ML PROFESSIONAL-DATA-ENGINEER Practice Questions (Page 4)

Part of the Preparing and using data for analysis domain, which accounts for ~15% of the PROFESSIONAL-DATA-ENGINEER exam. Google Cloud does not publish an official question count, but from its 120-minute exam (~50–80 total, ~8–12 in this domain), expect 3–4 from this objective — we provide 23 practice questions to prepare you well beyond it. (estimate)

23questions here
5free pages
5concepts
~15%of the exam

Questions 16–20

  1. 16foundation · easy

    In the context of retrieval-augmented generation (RAG), what is the purpose of chunking documents before generating embeddings?

    Select an answer first
  2. 17expert · hard

    A telecom company is developing a churn prediction model using BigQuery ML. The dataset includes a `call_duration` feature with many outliers, a `plan_type` categorical feature, and a `customer_tenure` feature. The data engineer needs to prepare the data for both training and batch prediction. They also need to ensure that the model is not overly influenced by extreme values. Which approach best addresses these requirements?

    Select an answer first
  3. 18application · medium

    A social media analytics company wants to generate embeddings for user posts to detect trending topics. They have a mix of short tweets and long blog posts. The embedding model they use has a maximum token limit of 512. What is the best approach to generate embeddings for this mixed-length content?

    Select an answer first
  4. 19application · medium

    A legal tech company is building a RAG system to answer questions from a large corpus of court rulings. They have already chunked the documents and generated embeddings. Now they need to index the embeddings in a vector store. Which additional step is essential to ensure high-quality retrieval?

    Select an answer first
  5. 20expert · hard

    A music streaming company wants to build a recommendation system using embeddings of song audio features. They have millions of songs and plan to use a pre-trained model to generate embeddings. The embeddings will be stored in a vector database for similarity search. The data engineering team needs to decide on a storage and indexing strategy that balances query latency and cost. Which approach is most appropriate?

    Select an answer first
Finished these 5 questions?

Review the revealed explanations, or continue through the curriculum.

Free Basic Practice is a study aid with revealable answers — not a scored exam. Examers.io is independent and not affiliated with or endorsed by Google Cloud. “PROFESSIONAL-DATA-ENGINEER” is a trademark of its owner, used for identification only.