# What is AI Red Teaming?

> AI red teaming is an authorized, structured adversarial evaluation of an AI system.

- Canonical URL: https://yellowcube.eu/glossary/ai-red-teaming/
- Publisher: Yellow Cube
- Language: en
- Contact: hello@yellowcube.eu

## Content

Testers adopt relevant attacker or misuse perspectives and attempt to produce unacceptable behavior, bypass controls, expose data, abuse connected capabilities or reveal weaknesses in the model, application and operating process. The target may include prompts, retrieval, model behavior, APIs, identities, tools, infrastructure and human oversight — not merely a chatbot conversation.

A useful engagement begins with a threat model, objectives, rules of engagement and measurable success criteria. Testers record reproducible evidence and the conditions that made each result possible, while owners prioritize findings, improve controls and retest. Methods may include manual exploration, automated test generation, known attack techniques and scenario-based exercises, but results need expert interpretation in the system’s actual business context.

### Key points

- **Scope and authorization:** Define the systems, data, users, providers, integrations, prohibited actions, escalation contacts and safeguards for availability, privacy and third parties.
- **Representative testing:** Cover direct and indirect prompt attacks, information disclosure, unsafe tool use, authorization failures, supply-chain assumptions and relevant model-level or societal harms.
- **Useful outputs:** Preserve prompts, inputs, model and application versions, responses, tool traces, expected behavior, impact, repeatability and proposed control improvements.
- **Operating rhythm:** Test before important releases and after changes to models, prompts, retrieval sources, tools, permissions or deployment context; track findings through remediation and verification.
- **Important limitation:** AI red teaming can discover failures but cannot prove that a system is safe or secure. Coverage is finite, model behavior can vary, and a point-in-time result may not represent later versions, contexts or better-resourced attackers.

### Related terms

[Red team](<https://yellowcube.eu/glossary/red-team/>) · [Penetration testing](<https://yellowcube.eu/glossary/penetration-testing/>) · [Threat modeling](<https://yellowcube.eu/glossary/threat-modeling/>) · [Adversarial machine learning](<https://yellowcube.eu/glossary/adversarial-machine-learning/>) · [Prompt injection](<https://yellowcube.eu/glossary/prompt-injection/>) · [AI security](<https://yellowcube.eu/glossary/ai-security/>)

### Sources

[OWASP GenAI Red Teaming Guide](https://genai.owasp.org/resource/genai-red-teaming-guide/) · [NIST AI 600-1, Generative AI Profile](https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence) · [MITRE ATLAS](https://atlas.mitre.org/)

## Attribution and scope

This Markdown representation is generated from the same approved content records as the canonical HTML page. Cite the canonical URL above when referencing this material. Product and service descriptions are informational; confirm project-specific requirements with Yellow Cube.

