What Is GPT-Red? OpenAI Built an AI Hacker to Attack Its Own Models

What Is GPT-Red? OpenAI Built an AI Hacker to Attack Its Own Models

TL;DR: GPT-Red in 30 Seconds GPT-Red is an automated red-teaming model OpenAI trained to break its own models; it is not public and never will be. It was trained with self-play: attacker GPT-Red and defender models played against each other, both getting stronger over time. In scenarios where human red-teamers succeeded only 13% of the time, GPT-Red hit 84%. GPT-5.6 was trained against GPT-Red, making it OpenAI’s most robust model against prompt injections (only a 0.05% failure rate). OpenAI just announced something unusual: not a model for users, but an “AI hacker” trained to break its own models. Meet GPT-Red! According to the research publication “Unlocking Self-Improvement for Robustness” released today, GPT-Red automates the work of human security testers (red teams) and does it far better than humans. ...

July 15, 2026 ·  Updated: July 15, 2026 ·  5 min ·  1020 words
Qwen3: Hybrid Thinking in 119 Languages

Qwen3: Hybrid Thinking in 119 Languages

Hello Qwen3! A New Era in AI 🚀 The latest member of the Qwen family, Qwen3, makes a bold entry into the world of large language models (LLMs). Officially announced on May 4, 2025, it features hybrid reasoning modes, impressive multilingual support, and enhanced agent capabilities. The Qwen3 series appeals to a wide audience by offering both MoE (Mixture of Experts) and dense model options. The flagship model, Qwen3-235B-A22B, delivers competitive performance in coding, mathematics, and general capabilities, rivaling top models like DeepSeek-R1, o1, o3-mini, Grok-3, and Gemini-2.5-Pro. The smaller MoE model, Qwen3-30B-A3B, outperforms QwQ-32B with 10x fewer active parameters, and even a small model like Qwen3-4B can match the performance of Qwen2.5-72B-Instruct. ...

May 4, 2025 ·  Updated: May 4, 2025 ·  6 min ·  1180 words