Examers.io
ExamsOrganizationsHow it worksPricingHelp & FAQ
ISTQB logo

Certified Tester Testing with Generative AI

Domain 1Objective 3

Multimodal LLMs and Vision-Language Models CT-GENAI Practice Questions (Page 1)

Part of the Introduction to Generative AI for Software Testing domain, which makes up ~12% of our current practice bank. ISTQB does not publish an official question count, but from its 60-minute exam (~25–40 total, ~3–5 in this domain), expect 1–1 from this objective — we provide 22 practice questions to prepare you well beyond it. (estimate)

22questions here
5free pages
5concepts

Questions 1–5

  1. 1expert · hard

    A team is building an automated visual regression testing system. They have a pre-trained vision-language model (VLM) that performs well on general image understanding. However, it often fails to detect subtle, application-specific UI changes (e.g., a button's color changing from #FF0000 to #FE0001). The team has a large, labeled dataset of historical UI screenshots and defect reports. What is the most effective strategy to improve the model's performance on this specific task?

    Select an answer first
  2. 2foundation · easy

    Which of the following best describes how a multimodal LLM processes input data?

    Select an answer first
  3. 3application · medium

    A QA team wants to automate the verification of a mobile banking app's login screen across different device resolutions. They have a set of screenshots and a text specification of the expected layout. Which approach best leverages a multimodal LLM for this task?

    Select an answer first
  4. 4application · medium

    A company wants to deploy a multimodal LLM to analyze thousands of UI screenshots nightly. The team is concerned about the cost of running this workload. What is the most significant cost factor they need to consider?

    Select an answer first
  5. 5application · medium

    A test automation engineer is exploring ways to automate the testing of a voice-controlled smart home app. They need a model that can understand the user's spoken command (audio) and the resulting state of the app's interface (image) to verify the correct action was taken. Which type of model is best suited for this end-to-end verification?

    Select an answer first
Finished these 5 questions?

Review the revealed explanations, or continue through the curriculum.

Free Basic Practice is a study aid with revealable answers — not a scored exam. Examers.io is independent and not affiliated with or endorsed by ISTQB. “CT-GENAI” is a trademark of its owner, used for identification only.