
Dell Data Science Optimize
Domain 3Objective 2
Text Preprocessing DATA-SCIENCE-OPTIMIZE Practice Questions (Page 5)
Part of the Natural Language Processing (NLP) domain, which accounts for 20% of the DATA-SCIENCE-OPTIMIZE exam.
44questions here
9free pages
15concepts
20%of the exam
Questions 21–25
- 21
A financial news classifier must distinguish between 'Q3 earnings up 5%' and 'Q3 earnings up five percent'. The model uses a vocabulary of the 10,000 most frequent tokens. How should numbers be handled to maximize classification accuracy?
Select an answer first - 22
Why is lowercasing commonly applied as a text preprocessing step?
Select an answer first - 23
Which tool is commonly used for lemmatization in Python's NLTK library?
Select an answer first - 24
A team is building a sentiment model for product reviews. They notice that 'Great' and 'great' are treated as different tokens, increasing vocabulary size. However, they also need to preserve proper nouns like 'Apple' (the company) for entity recognition. What is the best approach?
Select an answer first - 25
Which of the following is an example of a non-alphanumeric symbol that might be removed during preprocessing?
Select an answer first
Finished these 5 questions?
Review the revealed explanations, or continue through the curriculum.
Free Basic Practice is a study aid with revealable answers — not a scored exam. Examers.io is independent and not affiliated with or endorsed by Dell Technologies. “DATA-SCIENCE-OPTIMIZE” is a trademark of its owner, used for identification only.