×

Trending Technologies MCQs

AI Red Teaming MCQs (Multiple-Choice Questions)

Practice AI Red Teaming MCQs to test your knowledge of techniques used to identify vulnerabilities, security weaknesses, and potential misuse in artificial intelligence systems. These questions cover adversarial testing, prompt injection, jailbreaks, data leakage, model robustness, and responsible AI security assessment. They are useful for students, AI developers, cybersecurity professionals, and researchers preparing for technical interviews or strengthening their understanding of AI security. The set includes foundational concepts as well as practical testing scenarios.

AI Red Teaming MCQs

These AI Red Teaming multiple-choice questions cover important concepts such as adversarial attacks, large language model vulnerabilities, indirect prompt injection, sensitive information exposure, harmful output generation, and security evaluation methods. This set combines conceptual, technical, and scenario-based questions to help test your understanding of how AI systems are challenged and assessed before deployment.

AI Red Teaming MCQs cover threat modeling, attack surfaces, test-case design, tool-use security, multimodal risks, mitigation strategies, and post-assessment reporting. Each question includes an answer and explanation to help clarify the security principles and testing practices involved in evaluating AI systems.

List of AI Red Teaming MCQs

The following AI Red Teaming multiple-choice questions cover red team objectives, adversarial prompting, jailbreak testing, data extraction risks, retrieval-augmented generation (RAG) security, agent vulnerabilities, evaluation metrics, and real-world AI security scenarios.

1. What is the primary objective of AI red teaming?

  1. To increase the number of model parameters
  2. To identify weaknesses and risks by deliberately testing an AI system
  3. To eliminate the need for model evaluation
  4. To guarantee that a model never produces an incorrect answer

Answer: B) To identify weaknesses and risks by deliberately testing an AI system

Explanation:

AI red teaming involves testing an AI system under challenging or adversarial conditions to discover security vulnerabilities, unsafe behaviors, privacy risks, and failure modes before or after deployment.

2. Which activity is an example of AI red teaming?

  1. Measuring only the model's average response time
  2. Changing a website's color scheme
  3. Testing whether a chatbot reveals confidential information when given adversarial prompts
  4. Increasing the size of a training dataset without evaluating it

Answer: C) Testing whether a chatbot reveals confidential information when given adversarial prompts

Explanation:

Red teamers create carefully designed inputs to determine whether a system violates its security, privacy, or safety requirements. Attempts to expose confidential information are a common example.

3. What is a jailbreak attack against a large language model?

  1. An attempt to bypass the model's intended safety restrictions through crafted inputs
  2. A method for compressing model weights
  3. A process for improving network bandwidth
  4. A technique for converting text into numerical tokens only

Answer: A) An attempt to bypass the model's intended safety restrictions through crafted inputs

Explanation:

A jailbreak attempts to make a model disregard its intended restrictions or produce disallowed outputs. Red teamers test for such failures to evaluate the strength of the model's safeguards.

4. What does prompt injection attempt to manipulate?

  1. The physical temperature of a server
  2. The number of layers in a neural network
  3. The operating system's display resolution
  4. The model's interpretation of instructions by introducing conflicting or malicious content

Answer: D) The model's interpretation of instructions by introducing conflicting or malicious content

Explanation:

Prompt injection places instructions in user inputs or external content to influence the model's behavior. It can attempt to override intended directions, reveal data, or trigger unauthorized actions.

5. Why is threat modeling important before an AI red team exercise?

  1. It ensures that all attacks will succeed
  2. It identifies important assets, possible attackers, attack paths, and security objectives
  3. It removes the need for permission to conduct testing
  4. It guarantees that the AI model is unbiased

Answer: B) It identifies important assets, possible attackers, attack paths, and security objectives

Explanation:

Threat modeling helps a team understand what needs protection, who might attempt an attack, and how weaknesses could be exploited. It makes testing more focused and relevant to the system's actual risks.

6. Which of the following is an example of sensitive information exposure in an AI system?

  1. The chatbot uses a shorter sentence
  2. The model changes its response formatting
  3. The assistant reveals a private API key that was accidentally included in its accessible context
  4. The application displays a loading animation

Answer: C) The assistant reveals a private API key that was accidentally included in its accessible context

Explanation:

Sensitive information exposure occurs when an AI system discloses secrets, personal information, credentials, or confidential records to an unauthorized party. Red team exercises help identify paths that may lead to such disclosure.

7. What is the purpose of a red team scope document?

  1. To define authorized targets, testing boundaries, objectives, and restrictions
  2. To publish all discovered credentials publicly
  3. To guarantee that no defects exist
  4. To replace the need for test evidence

Answer: A) To define authorized targets, testing boundaries, objectives, and restrictions

Explanation:

A scope document establishes which systems may be tested, what techniques are permitted, what data must be protected, and when testing must stop. It reduces operational and legal risks.

8. What is an adversarial example in machine learning?

  1. A sample that always improves model accuracy
  2. A duplicate record used for database backup
  3. A model checkpoint stored after training
  4. An input deliberately modified to cause an incorrect or unexpected model prediction

Answer: D) An input deliberately modified to cause an incorrect or unexpected model prediction

Explanation:

An adversarial example is crafted to exploit a model's weaknesses. Small changes to an image, for example, may cause a classifier to assign an incorrect label while the image remains recognizable to a person.

9. Which practice makes an AI red team assessment more reproducible?

  1. Keeping attack prompts undocumented
  2. Recording test inputs, model versions, configurations, outputs, and evaluation criteria
  3. Changing the testing objective after every failed attempt
  4. Deleting all evidence immediately after testing

Answer: B) Recording test inputs, model versions, configurations, outputs, and evaluation criteria

Explanation:

Detailed records allow teams to reproduce findings, compare results across model versions, verify fixes, and understand the conditions under which a vulnerability occurred.

10. How does indirect prompt injection differ from direct prompt injection?

  1. Indirect injection can only affect image classification systems
  2. Direct injection requires a physical security breach
  3. Indirect injection places malicious instructions in external content the AI processes
  4. There is no difference between the two

Answer: C) Indirect injection places malicious instructions in external content the AI processes

Explanation:

Direct prompt injection is typically supplied through a user's input. Indirect prompt injection may be embedded in webpages, documents, emails, or retrieved passages that an AI system reads while completing a task.

11. What is the main security concern when an AI agent can execute external tools?

  1. The agent may perform unauthorized actions if tool access and instructions are not properly controlled
  2. The model automatically becomes more accurate
  3. The agent can no longer process text
  4. Tool access eliminates all prompt injection risks

Answer: A) The agent may perform unauthorized actions if tool access and instructions are not properly controlled

Explanation:

AI agents can interact with databases, files, APIs, and other services. Excessive permissions or weak approval controls may allow manipulated instructions to trigger unintended actions.

12. Which approach is most appropriate for testing a model's resistance to jailbreaks?

  1. Testing only ordinary greetings
  2. Checking whether the model returns answers quickly
  3. Disabling all logging before the exercise
  4. Using a documented set of authorized adversarial prompts and evaluating outputs against defined safety criteria

Answer: D) Using a documented set of authorized adversarial prompts and evaluating outputs against defined safety criteria

Explanation:

A structured jailbreak evaluation tests a range of attack patterns under controlled conditions. Defined criteria help determine whether responses violate the system's safety requirements.

13. What is the purpose of a baseline evaluation in AI red teaming?

  1. To remove all existing model safeguards
  2. To establish normal system behavior for comparison with adversarial test results
  3. To avoid testing the deployed configuration
  4. To guarantee that every model version behaves identically

Answer: B) To establish normal system behavior for comparison with adversarial test results

Explanation:

A baseline shows how a system performs under ordinary conditions. Comparing it with adversarial results helps distinguish expected behavior from security failures or regressions.

14. Which risk is particularly relevant to a retrieval-augmented generation (RAG) system?

  1. The monitor displays the wrong brightness
  2. The keyboard layout changes unexpectedly
  3. Unauthorized or malicious retrieved content influences the generated response
  4. The application uses too many heading tags

Answer: C) Unauthorized or malicious retrieved content influences the generated response

Explanation:

RAG systems retrieve external documents and provide them as context to a language model. Poor access controls, untrusted content, or malicious instructions in retrieved documents can cause data exposure or manipulated answers.

15. Why should red teamers test an AI system with different user roles?

  1. To verify that permissions and data access remain appropriate for each role
  2. To increase the model's vocabulary size
  3. To avoid testing authorization rules
  4. To make all users administrators

Answer: A) To verify that permissions and data access remain appropriate for each role

Explanation:

Role-based testing can reveal whether a user can access another user's records, administrative functions, or restricted documents. It checks the interaction between AI behavior and application authorization.

16. What does a false positive mean in an AI security evaluation?

  1. A real attack succeeds without detection
  2. The model has no security vulnerabilities
  3. A vulnerability is always reproduced in production
  4. A test incorrectly identifies safe behavior as a security failure

Answer: D) A test incorrectly identifies safe behavior as a security failure

Explanation:

A false positive occurs when an evaluation flags behavior that does not actually violate the defined security requirement. Human review and clear criteria help reduce inaccurate findings.

17. What is a false negative in AI red teaming?

  1. A harmless response is correctly classified as safe
  2. A genuine security failure is missed by the evaluation process
  3. A test detects every vulnerability
  4. A model refuses an ordinary request

Answer: B) A genuine security failure is missed by the evaluation process

Explanation:

A false negative occurs when a real vulnerability is not detected or is incorrectly classified as acceptable. It can leave a system exposed despite apparently successful testing.

18. Which measure is useful when evaluating whether an AI model leaks private training data?

  1. Checking the user interface's font size
  2. Measuring the number of buttons in the application
  3. Testing authorized data-exposure scenarios and assessing whether protected information can be recovered
  4. Counting the number of words in the system prompt only

Answer: C) Testing authorized data-exposure scenarios and assessing whether protected information can be recovered

Explanation:

Privacy-focused testing evaluates whether a model reveals information that should remain confidential. Results should be interpreted carefully because a disclosure may originate from training data, prompts, retrieval sources, or application storage.

19. What is the main purpose of a human reviewer in an AI red team assessment?

  1. To assess context, severity, and whether model behavior actually violates the evaluation criteria
  2. To make every automated test unnecessary
  3. To change all model outputs manually
  4. To guarantee that the model cannot be attacked

Answer: A) To assess context, severity, and whether model behavior actually violates the evaluation criteria

Explanation:

Human reviewers can interpret ambiguous outputs, recognize contextual risks, and validate automated findings. Human judgment is especially valuable for nuanced safety, privacy, and policy assessments.

20. Why is it important to test AI systems after a model or prompt update?

  1. Updates always eliminate existing vulnerabilities
  2. Testing is only necessary before deployment
  3. Updated models cannot introduce new behaviors
  4. An update may fix one weakness while introducing another

Answer: D) An update may fix one weakness while introducing another

Explanation:

Changes to model weights, system prompts, tools, retrieval sources, or safety filters can alter behavior. Regression testing helps verify previous fixes and detect newly introduced vulnerabilities.

21. Which of the following best describes AI model robustness?

  1. The model's ability to store unlimited data
  2. The model's ability to maintain acceptable behavior under relevant variations and adversarial inputs
  3. The number of users who visit the application
  4. The size of the model's user interface

Answer: B) The model's ability to maintain acceptable behavior under relevant variations and adversarial inputs

Explanation:

Robustness describes how reliably a model performs when inputs differ from ordinary examples or contain deliberate perturbations. Red team testing helps expose conditions under which performance or safety degrades.

22. What should a red teamer do if a test unexpectedly exposes a real production secret?

  1. Publish the secret to demonstrate the severity
  2. Continue extracting similar secrets without authorization
  3. Stop unnecessary access, protect the evidence, and follow the agreed incident-reporting process
  4. Send the secret to unrelated users for verification

Answer: C) Stop unnecessary access, protect the evidence, and follow the agreed incident-reporting process

Explanation:

Discovering a secret requires careful handling. The tester should avoid further exposure, preserve only necessary evidence, restrict access, and notify the designated security contacts according to the engagement rules.

23. What is the role of a system prompt in an LLM application?

  1. It provides high-level instructions that help shape the model's behavior
  2. It encrypts every database automatically
  3. It guarantees that user input is trustworthy
  4. It prevents all external content from influencing the model

Answer: A) It provides high-level instructions that help shape the model's behavior

Explanation:

A system prompt communicates intended behavior, constraints, and task requirements. However, it is not a substitute for access controls, input handling, output validation, or other security mechanisms.

24. Why should an AI application avoid relying only on prompt instructions for access control?

  1. Prompts cannot contain text
  2. Access control is unrelated to AI security
  3. Prompt instructions always override application code
  4. A model may misinterpret or disregard instructions, so permissions must also be enforced by trusted application logic

Answer: D) A model may misinterpret or disregard instructions, so permissions must also be enforced by trusted application logic

Explanation:

Natural-language instructions are not a reliable security boundary. Sensitive operations should enforce authentication, authorization, and data-access rules in trusted code outside the model.

25. Which scenario demonstrates a potential AI supply-chain risk?

  1. A user selects a larger font
  2. An application integrates a compromised model, plugin, or external AI dependency
  3. A model answers a routine arithmetic question correctly
  4. A developer documents the test environment

Answer: B) An application integrates a compromised model, plugin, or external AI dependency

Explanation:

AI supply-chain risks can arise from untrusted models, datasets, packages, plugins, and service providers. Red teamers assess whether these dependencies can introduce malicious behavior or undermine system security.

26. What is the primary purpose of testing denial-of-service risks in an AI application?

  1. To improve the visual appearance of responses
  2. To make every request require manual approval
  3. To determine whether resource exhaustion or excessive requests can disrupt service or create unreasonable costs
  4. To remove rate limits from the application

Answer: C) To determine whether resource exhaustion or excessive requests can disrupt service or create unreasonable costs

Explanation:

AI workloads may consume significant compute resources. Testing rate limits, request size restrictions, quotas, and timeouts can help reveal ways an attacker might exhaust capacity or drive up usage costs.

27. What does a successful mitigation validation demonstrate?

  1. The original attack no longer causes the documented failure under the tested conditions
  2. The model is permanently immune to all possible attacks
  3. The application no longer requires monitoring
  4. All security issues across the organization have been resolved

Answer: A) The original attack no longer causes the documented failure under the tested conditions

Explanation:

Mitigation validation repeats the relevant tests after a fix is applied. A passing result provides evidence that the specific issue was addressed under the tested conditions, but broader testing is still needed.

28. How can a red team assess whether a chatbot follows data-minimization principles?

  1. By measuring its response font size
  2. By checking only the number of generated tokens
  3. By verifying whether it uses the newest model architecture
  4. By testing whether it requests, retrieves, or discloses personal data beyond what the task requires

Answer: D) By testing whether it requests, retrieves, or discloses personal data beyond what the task requires

Explanation:

Data minimization limits the collection and use of personal information to what is necessary for a defined purpose. Testing can reveal unnecessary requests, excessive retrieval, or inappropriate disclosure.

29. Which technique can help discover unexpected harmful behaviors in a model?

  1. Testing one fixed prompt repeatedly and ignoring all other inputs
  2. Systematically varying prompts, contexts, roles, and input formats while recording outcomes
  3. Removing every evaluation criterion
  4. Testing only the model's first response after deployment

Answer: B) Systematically varying prompts, contexts, roles, and input formats while recording outcomes

Explanation:

Structured variation explores different ways an input or context may affect model behavior. It can reveal weaknesses that are not apparent from a small set of ordinary prompts.

30. What is an AI red team finding's severity rating intended to communicate?

  1. The total number of tokens used during testing
  2. The developer's personal opinion of the model
  3. The relative impact and risk of the discovered weakness, considering relevant context
  4. The number of people who wrote the report

Answer: C) The relative impact and risk of the discovered weakness, considering relevant context

Explanation:

Severity ratings help teams prioritize remediation. Useful factors include the sensitivity of affected assets, exploitability, required access, potential impact, and existing mitigations.

31. Why should red teamers evaluate multimodal AI systems using different input types?

  1. Images, audio, video, and text may introduce different attack surfaces and failure modes
  2. Multimodal systems only process numerical data
  3. Different input types cannot affect model behavior
  4. Testing one image proves that every modality is secure

Answer: A) Images, audio, video, and text may introduce different attack surfaces and failure modes

Explanation:

Multimodal systems process information through different pathways. An attack embedded in an image, spoken instruction, or video frame may expose weaknesses that ordinary text-only tests do not detect.

32. Which security control is most useful for limiting the impact of a compromised AI agent?

  1. Giving the agent permanent administrator access
  2. Allowing the agent to execute every tool without review
  3. Storing credentials in prompts so the model can retrieve them
  4. Applying least-privilege permissions and requiring approval for high-impact actions

Answer: D) Applying least-privilege permissions and requiring approval for high-impact actions

Explanation:

Least privilege restricts an agent to the resources and operations needed for its task. Human approval or additional policy checks for sensitive actions can reduce the consequences of prompt injection or incorrect decisions.

33. What is the purpose of a canary token in an authorized AI security test?

  1. To improve model training accuracy
  2. To provide a controlled indicator that can reveal unexpected access or disclosure
  3. To increase the number of model layers
  4. To replace all application authentication mechanisms

Answer: B) To provide a controlled indicator that can reveal unexpected access or disclosure

Explanation:

A canary token is a monitored, non-sensitive marker placed in a controlled test environment. If it appears in an unexpected output or triggers a monitored event, it can help reveal an exposure path.

34. Why is it risky to use untrusted AI-generated text directly in database queries or shell commands?

  1. Generated text cannot contain punctuation
  2. AI-generated text is always encrypted
  3. It may contain unsafe content that leads to injection or unintended command execution if handled incorrectly
  4. Database systems do not support text input

Answer: C) It may contain unsafe content that leads to injection or unintended command execution if handled incorrectly

Explanation:

Model output should be treated as untrusted input. Parameterized queries, safe APIs, strict validation, and command restrictions help prevent generated text from being interpreted as executable instructions.

35. Which statement about AI red teaming and traditional penetration testing is correct?

  1. AI red teaming can examine model-specific behaviors in addition to conventional application and infrastructure weaknesses
  2. AI red teaming never involves application security
  3. Traditional penetration testing automatically covers every AI-specific risk
  4. AI red teaming only tests physical server access

Answer: A) AI red teaming can examine model-specific behaviors in addition to conventional application and infrastructure weaknesses

Explanation:

AI red teaming may test prompt injection, unsafe generation, model privacy, data poisoning, and agent behavior. It can complement conventional penetration testing, which remains important for the surrounding application and infrastructure.

36. What is data poisoning in the context of machine learning security?

  1. Encrypting a training dataset with a strong algorithm
  2. Removing duplicate records from a dataset
  3. Reducing the size of a model checkpoint
  4. Manipulating training or fine-tuning data to influence model behavior in an unintended way

Answer: D) Manipulating training or fine-tuning data to influence model behavior in an unintended way

Explanation:

Data poisoning introduces malicious or misleading examples into a training process. Depending on the attack, it may degrade model performance, introduce targeted misbehavior, or create a backdoor.

37. What is a backdoor attack against a machine learning model?

  1. A method for backing up model files
  2. A hidden behavior that is triggered by a particular pattern or condition in the input
  3. A normal model update performed by an administrator
  4. A process for making every prediction transparent

Answer: B) A hidden behavior that is triggered by a particular pattern or condition in the input

Explanation:

A backdoored model may behave normally on most inputs but produce a targeted output when a trigger is present. Red team assessments can examine model provenance, suspicious training patterns, and trigger-dependent behavior.

38. Why should an AI red team test logging and monitoring mechanisms?

  1. To ensure that every prompt is published online
  2. To eliminate the need for access controls
  3. To determine whether important attacks and suspicious events are detected while sensitive logs remain protected
  4. To guarantee that all alerts are false positives

Answer: C) To determine whether important attacks and suspicious events are detected while sensitive logs remain protected

Explanation:

Effective monitoring helps security teams detect suspicious activity and investigate incidents. Testing should also check that logs do not unnecessarily store credentials, personal information, or confidential prompts.

39. What is the main risk of allowing a language model to make purchases or transfer funds without safeguards?

  1. Incorrect or manipulated instructions could cause unauthorized financial actions
  2. The model will always select the cheapest option
  3. The system will no longer need user authentication
  4. Financial transactions cannot be recorded

Answer: A) Incorrect or manipulated instructions could cause unauthorized financial actions

Explanation:

Financial actions have direct consequences. An AI system should use strict permissions, transaction limits, independent policy checks, and appropriate human confirmation rather than relying solely on model-generated decisions.

40. Which information should be included in a useful AI red team report?

  1. Only the name of the tested model
  2. Only a list of prompts with no results
  3. Only the total number of test cases
  4. The finding, affected component, reproduction steps, evidence, impact, severity, and recommended remediation

Answer: D) The finding, affected component, reproduction steps, evidence, impact, severity, and recommended remediation

Explanation:

A useful report enables developers and security teams to understand and reproduce the issue, assess its impact, and implement a fix. Clear evidence and actionable recommendations support verification and prioritization.

41. How can automated red teaming tools support AI security testing?

  1. By replacing every form of expert judgment
  2. By generating and evaluating many test cases consistently, subject to appropriate validation
  3. By proving that a model has no vulnerabilities after one run
  4. By disabling all safeguards in production

Answer: B) By generating and evaluating many test cases consistently, subject to appropriate validation

Explanation:

Automation can expand test coverage, repeat known attack patterns, and compare model behavior. However, automated evaluators can make mistakes, so significant findings may require human review.

42. What does attack success rate measure in an AI red team exercise?

  1. The total number of developers assigned to the project
  2. The number of model parameters
  3. The proportion of evaluated attack attempts that meet the defined success criteria
  4. The average length of every prompt

Answer: C) The proportion of evaluated attack attempts that meet the defined success criteria

Explanation:

Attack success rate measures how often an attack achieves its specified objective within a test set. The result depends on how success is defined, how cases are sampled, and which system configuration is evaluated.

43. Why is it important to test an AI system's refusal behavior?

  1. To check whether it appropriately declines disallowed requests while still helping with legitimate ones
  2. To make it refuse every user request
  3. To prevent it from answering harmless questions
  4. To remove all safety policies from the system

Answer: A) To check whether it appropriately declines disallowed requests while still helping with legitimate ones

Explanation:

Refusal testing examines both under-refusal, where unsafe requests are answered, and over-refusal, where legitimate requests are unnecessarily blocked. Balanced evaluation helps improve safety and usability.

44. What is a key concern when testing an AI system that summarizes private documents?

  1. Whether the summary uses bullet points
  2. Whether the summary has a title
  3. Whether the model uses complete sentences
  4. Whether users can obtain summaries containing information from documents they are not authorized to access

Answer: D) Whether users can obtain summaries containing information from documents they are not authorized to access

Explanation:

Document summarization must preserve access boundaries. Red teamers should test whether retrieval, caching, conversation history, or summarization logic can expose restricted content across user or organizational boundaries.

45. Which action is most appropriate when an AI red team discovers a vulnerability that affects multiple customers?

  1. Ignore the issue until the next annual review
  2. Escalate it through the agreed incident process and coordinate remediation based on the risk
  3. Share affected customer data with all testers
  4. Continue exploiting the issue against customers without permission

Answer: B) Escalate it through the agreed incident process and coordinate remediation based on the risk

Explanation:

A vulnerability with broad customer impact may require urgent containment, notification, and remediation. The team should follow the agreed escalation procedures and minimize unnecessary access to customer data.

46. Why should an AI red team assess third-party plugins and connectors?

  1. Plugins cannot access external systems
  2. Third-party services are always more secure than the main application
  3. Plugins may introduce excessive permissions, weak validation, or new paths for data exposure and unintended actions
  4. Connector testing only affects interface design

Answer: C) Plugins may introduce excessive permissions, weak validation, or new paths for data exposure and unintended actions

Explanation:

Connectors extend an AI application's capabilities and attack surface. Testing should examine permission boundaries, input validation, output handling, credential management, and the security of external service interactions.

47. What is the purpose of a controlled red team environment?

  1. To reduce the chance that tests cause unintended harm while allowing realistic security evaluation
  2. To ensure that findings cannot be reproduced
  3. To expose production credentials to all participants
  4. To bypass the agreed testing scope

Answer: A) To reduce the chance that tests cause unintended harm while allowing realistic security evaluation

Explanation:

A controlled environment can use synthetic data, isolated services, limited credentials, and monitoring to support realistic tests without unnecessarily risking real users, sensitive information, or production operations.

48. How can a team prioritize AI red team findings when many vulnerabilities are discovered?

  1. Fix only the findings with the longest descriptions
  2. Address issues in the order they were discovered, regardless of risk
  3. Ignore issues that require coordination across teams
  4. Consider impact, exploitability, affected assets, exposure, and the effectiveness of existing safeguards

Answer: D) Consider impact, exploitability, affected assets, exposure, and the effectiveness of existing safeguards

Explanation:

Risk-based prioritization directs resources toward the issues most likely to cause significant harm. A severe weakness affecting sensitive data or critical actions may deserve attention before a low-impact usability issue.

49. Which statement best describes continuous AI red teaming?

  1. It means running one test before the system is built
  2. It involves periodically or continuously reassessing risks as models, data, tools, and threats change
  3. It requires publishing all test prompts to the public
  4. It removes the need for security updates

Answer: B) It involves periodically or continuously reassessing risks as models, data, tools, and threats change

Explanation:

AI systems evolve as models, integrations, prompts, and data sources change. Ongoing assessments and regression tests help identify new weaknesses and verify that previous mitigations continue to work.

50. A company deploys an AI assistant that retrieves internal documents and can send emails through an API. During a red team test, a retrieved document contains instructions to email confidential files to an external address. What is the best security design to reduce this risk?

  1. Allow the assistant to follow all retrieved instructions automatically
  2. Hide the document's instructions from the user interface but leave email permissions unrestricted
  3. Treat retrieved content as untrusted, enforce recipient and data-access policies in application code, and require approval for sensitive email actions
  4. Store all confidential documents in the system prompt

Answer: C) Treat retrieved content as untrusted, enforce recipient and data-access policies in application code, and require approval for sensitive email actions

Explanation:

This scenario combines indirect prompt injection, confidential data exposure, and unsafe tool use. Retrieved text should not be allowed to grant itself authority. Trusted application logic must enforce document permissions, restrict email recipients and attachments, and apply confirmation or approval controls to sensitive actions. Red teamers should then retest the attack and related variations to validate the mitigation.

Comments and Discussions!

Load comments ↻



Copyright © 2026 www.includehelp.com. All rights reserved.