Examers.io
ExamsOrganizationsHow it worksPricingHelp & FAQ
Databricks logo

DatabricksCertified Generative AI Engineer Associate

Domain 2Objective 6

Use Tools and Metrics to Evaluate Retrieval Performance GENERATIVE-AI-ENGINEER-ASSOCIATE Practice Questions (Page 3)

Part of the Section 2: Data Preparation domain, which makes up ~25% of our current practice bank. Databricks does not publish an official question count, but from its 90-minute exam (~35–60 total, ~9–15 in this domain), expect 1–2 from this objective — we provide 18 practice questions to prepare you well beyond it. (estimate)

18questions here
4free pages
4concepts

Questions 11–15

  1. 11foundation · easy

    In the context of retrieval evaluation, what is 'ground truth'?

    Select an answer first
  2. 12application · medium

    A team is using the `ragas` library to evaluate retrieval performance in a Databricks notebook. They have a dataset with queries, retrieved documents, and ground-truth relevance labels. They want to compute the `context_precision` and `context_recall` metrics. What is the correct way to use the library?

    Select an answer first
  3. 13expert · medium

    A team is evaluating a RAG pipeline for a financial compliance system. They have a ground-truth set of 200 queries with relevance judgments. The retriever currently returns 5 documents per query. They compute recall@5 = 0.9 and precision@5 = 0.6. The compliance team requires that the final answer must be based on the most relevant document, and they are concerned about the impact of irrelevant documents in the context. They want to improve the pipeline's ability to surface the single most relevant document at the top of the list. Which metric should they focus on improving?

    Select an answer first
  4. 14application · medium

    A RAG pipeline retrieves 10 documents for each query. A human evaluator labels each retrieved document as relevant or not. For a particular query, 4 of the 10 retrieved documents are relevant, and the system's ranked list places 2 relevant documents in the top 3 positions. The team wants a single metric that captures both how many relevant documents were retrieved and how highly they were ranked. Which metric should they use?

    Select an answer first
  5. 15application · medium

    A RAG pipeline for a customer support chatbot retrieves 5 documents per query. The team evaluates the retriever on a test set and finds that precision@5 is 0.6, but recall@5 is 0.4. They want to improve the retriever so that it retrieves more of the relevant documents, even if it means retrieving some irrelevant ones. Which change is most aligned with this goal?

    Select an answer first
Finished these 5 questions?

Review the revealed explanations, or continue through the curriculum.

Free Basic Practice is a study aid with revealable answers — not a scored exam. Examers.io is independent and not affiliated with or endorsed by Databricks. “GENERATIVE-AI-ENGINEER-ASSOCIATE” is a trademark of its owner, used for identification only.