Security & governance · Glossary term

What is Red Teaming?

A structured adversarial testing process in which authorized testers seek failures using documented objectives, threat assumptions, cases, and evidence.

Why does Red Teaming matter?

Ordinary quality tests rarely explore how a system behaves under manipulation, misuse, conflicting goals, or determined attempts to bypass controls.

Red Teaming in practice

Derive attacks from a threat model, run them in an isolated environment, record reproducible cases, remediate by layer, and convert confirmed failures into regression evals.

What is the common confusion about Red Teaming?

A list of jailbreak prompts is not a complete red-team program, and red teaming cannot prove the absence of unknown failures.

Learn Red Teaming in the course

No lesson links to this term yet. Search the course catalog for it.

  • Threat ModelA documented account of protected assets, trust boundaries, potential adversaries, assumed capabilities, attack paths, impacts, and…
  • GuardrailsSystem controls that constrain inputs, tool use, outputs, permissions, and escalation.
  • Prompt InjectionAn attack or failure mode in which untrusted content influences a model to disregard intended instructions, expose data, misuse tools, or…
  • Eval SetA versioned collection of inputs, expected properties, scoring rules, and metadata used to measure an AI system against a defined…
  • JailbreakAn adversarial input or interaction strategy intended to make a model produce behavior that its training or application controls are…

Sources

More terms in Security & governance

Open the Security & governance list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.