
Microsoft said MDASH with MAI-Cyber-1-Flash scored 96 percent on CyberGYM, the standard benchmark test. The ranking is 12 points higher than Anthropic’s Mythos and also better than Google Gemini and OpenAI GPT. Using the new MDASH costs half as much as the previous MDASH offer.
The second tool Microsoft announced on Monday is called Project Perception. It is also a collection of specialized artificial intelligence agents that perform red, blue and green team functions to find vulnerabilities, investigate them to determine their risks and take corrective actions accordingly. Microsoft said the platform selects the models to use based on the assigned task. Considerations that go into the decision include the effectiveness of the model and the final cost to the customer. Microsoft said the decisions were shaped by “ongoing research, benchmarking and evaluation across frontier and specialized models.”
Microsoft he said Project Perception is designed to perform 90 percent of the tasks at a lower cost than similar platforms from competitors. This means that customers can only turn to more expensive alternatives for the remaining 10 percent of tasks.
Microsoft said the new tools respond to a seismic shift in how organizations protect their networks from catastrophic attacks.
“As artificial intelligence accelerates the speed and scale of cyberattacks, defenders are being asked to secure increasingly complex digital environments with approaches built for a different era,” the company said. “Security teams are often forced to combine signals, contexts and risk insights into large amounts of data, making it difficult to keep up with emerging threats.”
Last week’s OpenAI incident, reminiscent of disturbing scenes from the most dystopian sci-fi novels, the tools currently in preview deserve a healthy dose of caution, which Microsoft has made no mention of. They must be carefully inspected and evaluated before being used in production. On the other hand, there are obvious risks in not adopting such tools. Balancing the risks of using artificial intelligence agents and the dangers of avoiding them is a task that does not yet have clear answers.





