Testers adopt relevant attacker or misuse perspectives and attempt to produce unacceptable behavior, bypass controls, expose data, abuse connected capabilities or reveal weaknesses in the model, application and operating process. The target may include prompts, retrieval, model behavior, APIs, identities, tools, infrastructure and human oversight — not merely a chatbot conversation.
A useful engagement begins with a threat model, objectives, rules of engagement and measurable success criteria. Testers record reproducible evidence and the conditions that made each result possible, while owners prioritize findings, improve controls and retest. Methods may include manual exploration, automated test generation, known attack techniques and scenario-based exercises, but results need expert interpretation in the system’s actual business context.
Key points
Scope and authorizationDefine the systems, data, users, providers, integrations, prohibited actions, escalation contacts and safeguards for availability, privacy and third parties.
Representative testingCover direct and indirect prompt attacks, information disclosure, unsafe tool use, authorization failures, supply-chain assumptions and relevant model-level or societal harms.
Useful outputsPreserve prompts, inputs, model and application versions, responses, tool traces, expected behavior, impact, repeatability and proposed control improvements.
Operating rhythmTest before important releases and after changes to models, prompts, retrieval sources, tools, permissions or deployment context; track findings through remediation and verification.
Important limitationAI red teaming can discover failures but cannot prove that a system is safe or secure. Coverage is finite, model behavior can vary, and a point-in-time result may not represent later versions, contexts or better-resourced attackers.