Examers.io
ExamsOrganizationsHow it worksPricingHelp & FAQ
NetApp logo

NetAppCertified AI Expert

Domain 5Objective 1

Determine How to Size Storage and Compute for Training and Inferencing Workloads CERTIFIED-AI-EXPERT Practice Questions (Page 3)

Part of the AI Common Challenges domain, which accounts for 22% of the CERTIFIED-AI-EXPERT exam. NetApp does not publish an official question count, but from its 90-minute exam (~35–60 total, ~8–13 in this domain), expect 1–2 from this objective — we provide 18 practice questions to prepare you well beyond it. (estimate)

18questions here
4free pages
5concepts
22%of the exam

Questions 11–15

  1. 11foundation · easy

    An inference application reads small model metadata files for every request. The application requires consistent sub-millisecond response times. Which storage performance metric is most critical to evaluate?

    Select an answer first
  2. 12foundation · easy

    A training job uses a model that requires 16 GB of GPU memory per sample. The batch size is 8. What is the minimum GPU memory needed just to hold the model and the batch?

    Select an answer first
  3. 13application · medium

    A company is deploying a real-time inference service for a fraud detection model. The model is a gradient-boosted trees ensemble with an average inference latency of 5 ms on a CPU. The service must handle a peak of 2,000 requests per second with a p99 latency target of 50 ms. The team is deciding between scaling out with more CPU nodes or moving to GPU inference. Which consideration is most important when deciding between CPU and GPU inference for this workload?

    Select an answer first
  4. 14application · medium

    A company is deploying an image recognition model that processes images uploaded by users. The model is a convolutional neural network that takes 20 ms to process an image on a GPU. The service must handle 200 requests per second. The team is deciding between using a single GPU node with 8 GPUs or multiple nodes with 2 GPUs each. What is the most important factor in determining the number of GPUs needed?

    Select an answer first
  5. 15foundation · easy

    A model's inference latency is 50 ms on a single GPU. The service-level agreement (SLA) requires a p95 latency of under 100 ms. What is the primary compute resource consideration for this workload?

    Select an answer first
Finished these 5 questions?

Review the revealed explanations, or continue through the curriculum.

Free Basic Practice is a study aid with revealable answers — not a scored exam. Examers.io is independent and not affiliated with or endorsed by NetApp. “CERTIFIED-AI-EXPERT” is a trademark of its owner, used for identification only.