Red Team AI

Noun · AI & Machine Learning

Definitions

  1. The practice of adversarially testing AI systems to uncover unsafe behavior, security weaknesses, misuse pathways, or policy failures before deployment. Red teaming is a common safety practice for high-risk AI features.

    In plain English: Adversarial testing of an AI system to find weaknesses before release.

    Example: "They ran a red team AI exercise to probe prompt injection, unsafe advice, and data leakage before expanding access."

Related Terms