
Dell Data Science Optimize
Domain 3Objective 2
Text Preprocessing DATA-SCIENCE-OPTIMIZE Practice Questions (Page 6)
Part of the Natural Language Processing (NLP) domain, which accounts for 20% of the DATA-SCIENCE-OPTIMIZE exam.
44questions here
9free pages
15concepts
20%of the exam
Questions 26–30
- 26
Why might you filter out rare words during text preprocessing?
Select an answer first - 27
A multilingual dataset contains text with inconsistent Unicode encodings, including full-width characters (e.g., 'A' vs 'A') and varying whitespace. The team needs consistent tokens for a cross-lingual model. What should they do?
Select an answer first - 28
A sentiment model for product reviews has a vocabulary of 50,000 tokens, but many appear only once in the training data. The team wants to reduce overfitting and improve generalization. What should they do?
Select an answer first - 29
A sentiment model for customer support chats must handle informal text like 'cant', 'wont', and 'im'. The team is debating whether to expand contractions before or after tokenization. The tokenizer splits on whitespace and punctuation. What is the correct approach?
Select an answer first - 30
A chatbot training dataset contains informal messages like 'idk', 'gonna', and 'don't'. The team wants to standardize these before training a language model. Which preprocessing step should be applied before tokenization?
Select an answer first
Finished these 5 questions?
Review the revealed explanations, or continue through the curriculum.
Free Basic Practice is a study aid with revealable answers — not a scored exam. Examers.io is independent and not affiliated with or endorsed by Dell Technologies. “DATA-SCIENCE-OPTIMIZE” is a trademark of its owner, used for identification only.