OpenAI built an attacker that breaks its own models 6× better than human experts — so the models you use are already hardened against it.
GPT-Red is an automated red-teaming system trained by self-play: it is rewarded for breaking defender models while the defenders train to resist. It succeeded in 84% of attack scenarios versus 13% for human testers, discovered a novel "fake chain-of-thought" attack class before human researchers did, and its findings cut prompt-injection failures 6-fold on OpenAI's hardest benchmark.
Ask every AI vendor in your stack how they red-team for prompt injection — adversarial testing is becoming table stakes, not a nice-to-have.
Prompt injection is exactly the risk logistics AI agents face in the wild: a malicious instruction hidden in an email, a rate confirmation, or a carrier document can hijack an agent that books, quotes, or pays. In one test, GPT-Red manipulated an autonomous vending agent into lowering prices and cancelling a customer order — swap "vending" for "tendering" and the lesson writes itself.
https://openai.com/index/unlocking-self-improvement-gpt-red/