
Google CloudProfessional Machine Learning Engineer
Domain 4Objective 2
Scaling Online Model Serving PROFESSIONAL-MACHINE-LEARNING-ENGINEER Practice Questions (Page 5)
Part of the Serving and scaling models domain, which accounts for ~20% of the PROFESSIONAL-MACHINE-LEARNING-ENGINEER exam.
37questions here
8free pages
14concepts
~20%of the exam
Questions 21–25
- 21
What is the purpose of model quantization?
Select an answer first - 22
A company serves a real-time object-detection model for autonomous vehicles. The model must run on a device in the vehicle with limited power and no internet connection. The team has a trained TensorFlow model. What should they do to deploy it?
Select an answer first - 23
Which monitoring metric is important to track for features served online?
Select an answer first - 24
What is a canary deployment in model serving?
Select an answer first - 25
A team deploys a real-time translation model on Vertex AI Endpoints. The model is CPU-based and currently handles 100 QPS with a p99 latency of 150ms. The team expects traffic to double to 200 QPS next month. They want to maintain the same latency SLA. What is the most cost-effective approach?
Select an answer first
Finished these 5 questions?
Review the revealed explanations, or continue through the curriculum.
Free Basic Practice is a study aid with revealable answers — not a scored exam. Examers.io is independent and not affiliated with or endorsed by Google Cloud. “PROFESSIONAL-MACHINE-LEARNING-ENGINEER” is a trademark of its owner, used for identification only.