
DatabricksCertified Generative AI Engineer Associate
Domain 2Objective 2
Filter Extraneous Content in Source Documents That Degrades Quality of a RAG Application GENERATIVE-AI-ENGINEER-ASSOCIATE Practice Questions (Page 4)
Part of the Section 2: Data Preparation domain, which makes up ~25% of our current practice bank. Databricks does not publish an official question count, but from its 90-minute exam (~35–60 total, ~9–15 in this domain), expect 1–2 from this objective — we provide 20 practice questions to prepare you well beyond it. (estimate)
20questions here
4free pages
4concepts
Questions 16–20
- 16
A retail company is building a RAG application over product reviews. The reviews contain a lot of noise, including 'This is a sponsored review', 'I received this product for free', and repetitive promotional text. The team wants to filter this noise to improve the quality of the RAG responses. Which approach is most appropriate?
Select an answer first - 17
After applying a new filtering step to a RAG pipeline, what is the MOST direct way to evaluate its impact on retrieval quality?
Select an answer first - 18
During document ingestion for a RAG pipeline, which of the following is an example of extraneous content that should be identified and removed?
Select an answer first - 19
A team is building a RAG application over a corpus of academic papers. They apply a filtering step that removes all 'Acknowledgements' sections, as these are considered noise. After evaluation, they find that retrieval precision has improved, but the answer generation for queries about 'funding sources' has degraded. The team is considering whether to revert the filter. What is the best way to evaluate the impact of this filter?
Select an answer first - 20
A healthcare company is building a RAG system over clinical trial protocols. The documents contain a 'Statistical Analysis Plan' section that includes complex formulas and tables, alongside narrative descriptions of the analysis methods. The team applies a filtering step that removes all lines containing mathematical symbols (e.g., '=', '+', 'σ') to clean the text. After deployment, they notice that retrieval for queries about 'how the primary endpoint was analyzed' has degraded significantly. What is the most likely cause?
Select an answer first
Finished these 5 questions?
Review the revealed explanations, or continue through the curriculum.
No more pagesBack to GENERATIVE-AI-ENGINEER-ASSOCIATE
Free Basic Practice is a study aid with revealable answers — not a scored exam. Examers.io is independent and not affiliated with or endorsed by Databricks. “GENERATIVE-AI-ENGINEER-ASSOCIATE” is a trademark of its owner, used for identification only.