Джейлбрейк ИИ: новый инструмент обходит защиту моделей Google, OpenAI, Anthropic и xAI
Wired протестировал новый инструмент для джейлбрейка на флагманских моделях Google, OpenAI, Anthropic и xAI. Итог: защиту части передовых моделей обойти тревожно легко, а устойчивость к взлому у разных компаний ощутимо различается — одни держат оборону лучше, другие заметно слабее.
AI-processed from Wired; edited by Hamidun News
Wired tested a new jailbreak tool — a way to bypass built-in safety restrictions — on the models of four of the largest AI companies: Google, OpenAI, Anthropic, and xAI. The correspondent's conclusion is blunt: the safeguards of some frontier models turn out to be alarmingly easy to strip away.
What Wired's experiment showed
A Wired correspondent watched as a new tool tried, one by one, to bypass the safety mechanisms of models from four frontier companies. These are the makers of Gemini (Google), GPT (OpenAI), Claude (Anthropic), and Grok (xAI) — essentially the entire top tier of the industry. According to the outlet's observations, the models behaved differently: some held up better, others noticeably worse, and it's exactly this spread that the author calls unexpected.
Key facts from the piece:
- The tool was tested against models from four companies: Google, OpenAI, Anthropic, and xAI
- The targets were the safeguards of flagship models — Gemini, GPT, Claude, and Grok
- Wired's main finding: the safeguards of some frontier models are alarmingly easy to bypass
- Resistance to jailbreaking varies noticeably between companies
"I watched as a new tool tried to circumvent the safeguards of four major frontier companies.
You might be surprised by how they performed," Wired's report states.
What a model jailbreak is
A jailbreak is a technique that makes a language model produce output it's not supposed to under its own rules: instructions for making weapons, malicious code, disinformation, or other prohibited content. Formally, a model has filters and rules, but a specially crafted prompt gets around them.
Classic methods include role-play scenarios ("imagine you're a villain from a movie"), disguising the request as an academic exercise, gradual escalation from harmless to prohibited, or context substitution. Newer tools automate this trial-and-error: instead of manually crafting phrasings, they run hundreds of variants through the model until one works. All four companies in the test publicly invest in red-teaming — adversarial testing of models before release — but jailbreaking remains one of the main classes of attacks on large language models.
Why this is dangerous
The danger is that a successful jailbreak turns a helpful assistant into a source of real harm. As long as a model just answers in a chat window, the cost of a failure is limited; but frontier models increasingly get access to tools, a browser, and actions taken on the user's behalf — and once that happens, stripped-away safeguards scale up the consequences. Unlike a typical software vulnerability, a jailbreak can't be closed with a single patch: a model isn't a line of code but learned behavior, and patching one loophole often opens another.
The spread of results across the four companies is a signal in its own right. It means that "frontier model safety" is not a single industry standard but a property of a specific product: some have higher barriers, while the very same tool strips others down. For a business embedding these models into its services, that's the difference between managed and unmanaged risk.
What this means
Wired's test is another reminder that the safeguards of top models aren't absolute, and their reliability varies noticeably from company to company. As AI assistants gain more rights and autonomy, resistance to jailbreaking will stop being an abstract topic for researchers and become a practical criterion for choosing a model.
Need AI working inside your business — not just in your newsfeed?
I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.