
Dell Data Science Optimize
Domain 3Objective 2
Text Preprocessing DATA-SCIENCE-OPTIMIZE Practice Questions (Page 8)
Part of the Natural Language Processing (NLP) domain, which accounts for 20% of the DATA-SCIENCE-OPTIMIZE exam.
44questions here
9free pages
15concepts
20%of the exam
Questions 36–40
- 36
A fraud-detection model for transaction descriptions must distinguish 'transfer $100' from 'transfer $10000'. The vocabulary is capped at 20,000 tokens, and numbers appear in many formats ('$100', '100.00', 'one hundred'). What is the best number-handling strategy?
Select an answer first - 37
A web-crawling pipeline collects forum posts that contain HTML tags, JavaScript code snippets, and URLs. The team wants to extract only the visible text for a topic classifier. They are considering two orders: (A) strip tags then remove URLs, or (B) remove URLs then strip tags. Which is correct and why?
Select an answer first - 38
Which of the following is an example of a tokenizer that splits text based on whitespace?
Select an answer first - 39
A team is training an LSTM for document classification. Documents vary in length from 10 to 2,000 tokens. The model requires fixed-length input. What is the most appropriate strategy?
Select an answer first - 40
A team is building a document retrieval system for legal contracts. They need to match queries like 'termination' with documents containing 'terminate' and 'terminated'. They also need to distinguish 'saw' (tool) from 'saw' (past tense of see) in context. Which approach best balances recall and precision?
Select an answer first
Finished these 5 questions?
Review the revealed explanations, or continue through the curriculum.
Free Basic Practice is a study aid with revealable answers — not a scored exam. Examers.io is independent and not affiliated with or endorsed by Dell Technologies. “DATA-SCIENCE-OPTIMIZE” is a trademark of its owner, used for identification only.