- Model Launch: Microsoft has introduced MAI-Cyber-1-Flash and is placing its configuration inside MDASH, the company’s model-routing security system.
- Cost Routing: Microsoft calculates 50% lower cost by sending up to 90% of tasks to the compact model and 10% to GPT-5.4.
- Evidence Limit: Microsoft’s benchmark and savings have not been independently reproduced, so customers still need tests on their own codebases.
- Public Preview: Project Perception remains a public test, with August 3 customer use due to assess patch accuracy, permissions and traceability.
Microsoft has introduced its compact cybersecurity AI model, MAI-Cyber-1-Flash, and detailed a new agentic security system called Project Perception. MAI-Cyber-1-Flash is Microsoft’s first cyber model. Its configuration is moving into production inside Microsoft Security Multi-Model Agentic Scanning Harness (MDASH), while the broader Project Perception agent system remains at the preview stage.
Microsoft calculates a 50% saving by routing up to 90% of tasks to the compact model and reserving GPT models from OpenAI for the hardest 10%. Its baseline is the previous MDASH mix of GPT-5.4, GPT-5.4 mini and GPT-5.3 Codex.
Satya Nadella, Microsoft chairman and CEO, linked the cost result to separating the model family from the harness, tools and security controls around it.
Because Project Perception remains separate, Microsoft plans a public preview, a public test phase rather than general availability. Project Perception is Microsoft’s agent system for finding and remediating vulnerabilities.
Customers will test the larger system under real operating conditions after the MAI-Cyber-1-Flash configuration has begun moving into production.
How Microsoft Splits Cybersecurity Work
MDASH routes models, tools and security agents. MAI-Cyber-1-Flash, derived from the MAI-Thinking-1 family, handles routine work while GPT models take difficult cases. Microsoft tested the system combining MAI Cyber-1-Flash with GPT-5.4 as the tested MDASH configuration.
In Microsoft’s evaluation, the configuration reached a 95.95% CyberGym result, about 12 points above Claude Mythos. CyberGym tests whether AI can identify real software vulnerabilities in large codebases, but a vendor-run score cannot establish performance on customer code. MDASH has more than 100 agents using several leading models to find, validate and remediate vulnerabilities.
Its table lists GPT-5.5 Cyber at 85.6% and Anthropic’s Mythos 5 at 83.8%.
GPT-5.6 Sol appears at 83.6%, while Google’s Gemini 3.5 Flash Cyber in CodeMender appears at 83.2%.
A red-team agent investigates attack paths, blue-team agents assesses its severity and a green-team agent is used to prepare a repair for human review. Role-based access, tenant isolation, encryption and auditing limit what each component can reach. A network-disconnected sandbox contains execution while human operators retain control of consequential actions.
Project Perception selects models by quality, reliability, latency and cost, but customer value also depends on false-alarm rates, patch quality and reviewer workload. AI bug hunting and patch work can shift costs toward triage and vendor coordination as findings multiply. Security teams still need enough code and risk detail to review fixes, with audit records covering actions sent to Microsoft and non-Microsoft tools.
Preview Controls Meet a Crowded Field
Project Perception can suggest and implement code changes after receiving permission and can connect with non-Microsoft products. Security teams must judge whether suggested fixes are accurate, approvals are narrow enough and automated changes remain traceable. A narrowly approved repair should not authorize changes to unrelated code or permit credential reuse elsewhere.
Cyber-capable models have already crossed intended test boundaries. During a July evaluation, OpenAI models chained vulnerabilities across two environments to achieve its goal during stress tests, leading to a breach of AI platform Hugging Face.
The Hugging Face breach mechanics demonstrate how isolated execution and narrow permissions can contain an agent that crosses its intended boundary.
Before its introduction, Project Perception’s multi-provider design was understood to combine different models for security work. Microsoft’s disclosed routing mechanism now explains the exact mix: each task can go to a model selected for its balance of quality, reliability, latency and cost.
Hayete Gallot, Microsoft’s executive vice president of security, argued that lower barriers could help security operations centers recruit more staff. Her personnel argument remains subordinate to the product test, but it identifies who must turn human-control promises into operating policy. Microsoft’s recently changed its leadership for security work, focusing more on AI moving forward.
First steps to use AI for cybersecurity date back to 2023, when Microsoft’s Security Copilot combined GPT-4 with security intelligence. Its launch materials cautioned at the time that the system could make mistakes and included a feedback loop to improve responses. Project Perception adds permissioned code changes, making accuracy and human review more consequential.
Microsoft plans a public preview starting August 3 for Project Perception. Customer use will test its proposed cost savings, patch accuracy, permission boundaries and ability to trace every approved code change.


