TechCrunch→ original

Ограничения ИИ от OpenAI и Anthropic мешают исследователям кибербезопасности

Издание TechCrunch опросило нескольких специалистов по наступательной кибербезопасности — тех, кто ищет неизвестные уязвимости и пишет инструменты для их эксплуатации. Главная жалоба: защитные ограничения (guardrails) моделей OpenAI и Anthropic всё чаще отказывают в помощи с задачами, которые для «белых» хакеров — рутина. Причина в том, что тот же запрос мог бы прийти и от злоумышленника, а модель не видит намерения.

AI-processed from TechCrunch; edited by Hamidun News
Ограничения ИИ от OpenAI и Anthropic мешают исследователям кибербезопасности
Source: TechCrunch. Collage: Hamidun News.
◐ Listen to article

TechCrunch published a piece on July 23, 2026, about how the guardrails of OpenAI's and Anthropic's language models are hindering the work of offensive cybersecurity researchers — specialists who search for as-yet-unknown vulnerabilities and develop tools to exploit them.

Why AI filters get in pentesters' way

OpenAI's models (ChatGPT) and Anthropic's (Claude) are increasingly refusing to help with offensive security tasks, even when it's a "white hat" specialist at the keyboard. According to TechCrunch, the outlet interviewed several researchers who professionally hunt for zero-day vulnerabilities and write exploits — and they all say the same thing: the filters trigger on a legitimate work task as if it were a hacking attempt.

The problem is that a request like "write code that exploits this specific vulnerability" looks identical to the model regardless of who sent it — a company's red team or an attacker. The model only sees the text of the request, not the intent behind it or the pentest contract backing it.

  • Outlet — TechCrunch, publication date — July 23, 2026
  • Models in focus — OpenAI (ChatGPT) and Anthropic (Claude)
  • Who was interviewed — offensive cybersecurity (offensive security) researchers
  • The core complaint — guardrails refuse to help with vulnerability research and exploit writing

Guardrails are safety rules built into a model that block potentially harmful output: instructions for building weapons, malicious code, or ways to bypass defenses. OpenAI and Anthropic position them as a key element of responsible AI, but it's precisely at the boundary with offensive security that the rules trigger too broadly.

What the dual-use dilemma is

Offensive cybersecurity is a classic example of dual-use technology: the same skills and tools that protect infrastructure can also break it. A red team impersonates an attacker to find holes before a real adversary does — and for that it needs exactly the same exploits, malware emulation, and defense-evasion techniques.

OpenAI's and Anthropic's guardrails are tuned to block precisely this class of request, because the labs can't tell a researcher from a criminal on the fly. As a result, as TechCrunch describes it, security professionals keep running into refusals in places where, just a year or two ago, the model would happily help — forcing them to work around the restrictions or switch to less regulated tools.

OpenAI's and

Anthropic's guardrails are getting in the way of offensive cybersecurity researchers' work — that's how the core of the problem is framed in TechCrunch's piece from July 23, 2026.

What this means for defense

Over-tightened filters don't just hit the attacking side — they slow down defense too. If a legitimate red team can't use a top-tier model for routine work, it loses speed, and the model itself loses usefulness for an entire professional segment. According to TechCrunch, some of those interviewed are switching to local open-weight models or narrowly specialized security tools that don't have corporate filters — at the cost of quality and convenience.

That creates a further imbalance: attackers already work with unrestricted models (open-weight builds, "uncensored" versions, underground services), while it's precisely those acting legally and under contract whose hands end up tied. Following TechCrunch's argument, OpenAI's and Anthropic's excessive caution risks punishing conscientious researchers without stopping the bad-faith ones.

What it means

The debate over guardrails in offensive security is a debate about where the line falls between "safe by default" and "useless for professionals." Until OpenAI and Anthropic learn to distinguish a legitimate pentest from attack preparation — for example, through organization verification or special modes for security teams — researchers will keep choosing the tools that don't block them.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Need AI working inside your business — not just in your newsfeed?

I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).

What do you think?
Loading comments…