Examers.io
ExamsOrganizationsHow it worksPricingHelp & FAQ
Databricks logo

DatabricksCertified Generative AI Engineer Associate

Domain 2Objective 1

Apply a Chunking Strategy for a Given Document Structure and Model Constraints GENERATIVE-AI-ENGINEER-ASSOCIATE Practice Questions (Page 1)

Part of the Section 2: Data Preparation domain, which makes up ~25% of our current practice bank. Databricks does not publish an official question count, but from its 90-minute exam (~35–60 total, ~9–15 in this domain), expect 1–2 from this objective — we provide 19 practice questions to prepare you well beyond it. (estimate)

19questions here
4free pages
5concepts

Questions 1–5

  1. 1application · medium

    A team has implemented a chunking strategy for a set of medical research papers. They are evaluating the quality of the chunks by checking if each chunk is self-contained and coherent. They find that some chunks contain only a table and its caption, while others contain only the text that references the table. What is the most likely problem with their current strategy?

    Select an answer first
  2. 2foundation · easy

    A generative AI engineer is selecting a chunk size for a document to be processed by a model with a 4,096-token context window. What is the primary constraint that determines the maximum chunk size?

    Select an answer first
  3. 3foundation · easy

    When analyzing a document's structure to inform chunking, which element is most likely to represent a natural semantic boundary?

    Select an answer first
  4. 4application · medium

    A company is building a chatbot to answer questions from its internal policy documents. The documents are a mix of short paragraphs and long, dense tables. The target LLM has a context window of 4096 tokens. The team wants to ensure that the most relevant information is retrieved for each query. What is the most important consideration when chunking the tables?

    Select an answer first
  5. 5application · medium

    A team is building a RAG system for a large corpus of legal documents. The documents are very long, often exceeding 10,000 tokens. The target LLM has a context window of 4096 tokens. They are using a recursive chunking strategy with a chunk size of 3000 tokens and an overlap of 200 tokens. What is the most important thing to verify about their chunks?

    Select an answer first
Finished these 5 questions?

Review the revealed explanations, or continue through the curriculum.

Free Basic Practice is a study aid with revealable answers — not a scored exam. Examers.io is independent and not affiliated with or endorsed by Databricks. “GENERATIVE-AI-ENGINEER-ASSOCIATE” is a trademark of its owner, used for identification only.