MarkTechPost→ original

Microsoft выпустила MAI-Cyber-1-Flash: 5B активных параметров и 95,95% на CyberGym

Microsoft AI выпустила MAI-Cyber-1-Flash — первую модель, заточенную под киберзащиту: 137 млрд параметров всего, 5 млрд активных, контекст 256к токенов. В связке со сканером MDASH она подняла результат на бенчмарке CyberGym до 95,95% против 88,45% в мае. Модель берёт на себя до 90% задач, а сложные 10% уходят на GPT-5.4 — заявленная экономия около 50%.

AI-processed from MarkTechPost; edited by Hamidun News
Microsoft выпустила MAI-Cyber-1-Flash: 5B активных параметров и 95,95% на CyberGym
Source: MarkTechPost. Collage: Hamidun News.
◐ Listen to article

Microsoft AI released MAI-Cyber-1-Flash on July 28, 2026 — its first model built specifically for cyber defense. Paired with the scanning harness MDASH, it pushed the CyberGym benchmark score to 95,95%, beating the industry's previous record.

What kind of model MAI-Cyber-1-Flash is

MAI-Cyber-1-Flash is a sparse Mixture-of-Experts transformer from Microsoft AI with 137 billion parameters in total, of which 5 billion are active, a 256-thousand-token context window, and text-only input and output. It is a specialized fine-tune of the MAI-Code-1-Flash model — a lightweight agentic model for code already built into GitHub Copilot and VS Code, from the MAI-Thinking-1 lineup. The model is not shipped as a separate API endpoint: it works only inside MDASH.

  • 137 billion parameters total, 5 billion active, sparse MoE architecture
  • 256-thousand-token context, input and output — text only
  • Specialized fine-tune of MAI-Code-1-Flash from the MAI-Thinking-1 lineup
  • 95,95% on CyberGym level 1 — versus 88,45% for MDASH in May 2026
  • The model handles up to 90% of MDASH's tasks, the tough 10% goes to GPT-5.4

Why the record is set by routing, not the model

The main result comes not from the model itself, but from the system around it. MDASH orchestrates more than 100 specialized agents through five stages: Prepare, Scan, Validate, Dedupe, and Prove. The CyberGym benchmark is a public set of 1507 real-world vulnerability-reproduction tasks collected from 188 OSS-Fuzz projects; in the level 1 configuration, the system is given the vulnerable source code and a general description.

MAI-Cyber-1-Flash handles up to 90% of all MDASH tasks, escalating the hardest 10% to GPT-5.4 — this delivers the claimed 50% cost savings versus the previous combination of GPT-5.4, 5.4 mini, and 5.3 codex. MDASH's previous maximum in May 2026 — 88,45% — was achieved using only publicly available models and was already the best public result, outperforming the nearest competitor's 83,1% by about five points.

"Replacing 80% of the models inside MDASH moved the harness from 88,4% to 95,95%," the release description from

Microsoft's research team states.

Why the model can't write exploits

MAI-Cyber-1-Flash's zeros on the ExploitGym benchmark (0/0/0 across the Kernel, Userspace, and Browser categories) are not a flaw but a deliberate Microsoft decision. According to the team, the model was trained to perform defensive tasks like bug patching, but not offensive ones like deploying malicious code.

In real-world work, the pairing proved itself convincing. In May 2026, work involving MDASH produced 16 CVEs in Windows' networking and authentication stack, including four critical remote code execution vulnerabilities. Retrospectively, the harness recovered 96% of 28 MSRC cases in the clfs.sys driver and 100% of 7 cases in tcpip.sys over a five-year window. MDASH itself was built by the Autonomous Code Security team, which includes members of Team Atlanta — winner of the DARPA AI Cyber Challenge.

What this means

Microsoft is showing that records in vulnerability discovery are set not by a single large model, but by smart orchestration of cheap specialized agents: a model with 5 billion active parameters carries 90% of the work and cuts costs in half, while the frontier model is brought in only for the hardest cases.

FAQ

What is MDASH?

MDASH is Microsoft's multi-model agentic harness for vulnerability discovery. It orchestrates more than 100 agents through five stages (Prepare, Scan, Validate, Dedupe, Prove) and scores 95,95% on the CyberGym benchmark.

Why does MAI-Cyber-1-Flash score zero on ExploitGym?

This is intentional. According to Microsoft, the model was trained only on defensive tasks — such as bug patching — and was not trained to write exploits or malicious code.

How much cheaper is the new pairing?

Microsoft claims savings of about 50% compared to the previous configuration of GPT-5.4, 5.4 mini, and 5.3 codex — because MAI-Cyber-1-Flash handles up to 90% of tasks, leaving only the hard 10% to GPT-5.4.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Need AI working inside your business — not just in your newsfeed?

I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).

What do you think?
Loading comments…