Wiz researchers found a single credential that could unlock every database running on Microsoft’s Azure Cosmos DB. They named it the Cosmos Master Key. And the vulnerability hunter that helped find it was not entirely human.
Microsoft has patched the flaw, disclosed on Thursday and dubbed CosmosEscape. It found no evidence anyone exploited it beyond Wiz’s own testing, and says customers need to do nothing, Reuters reported.
The scope was the alarming part. Cosmos DB is core Azure plumbing, and Microsoft services such as Teams, Entra ID and Copilot all store data in it. A working exploit could have listed every account in a region, then pulled the primary key for any of them. That grants full read and write access. The exposure reached Microsoft’s own internal databases, too.
Found with help from an AI
Buried in Wiz’s write-up is the detail that matters for where security is heading. The research “was assisted by an early version of Atlas,” the company’s AI vulnerability researcher. The same class of tool that broke into companies this month just helped find a master key to a flagship cloud.
Atlas is not one model. This week Wiz said it beat Anthropic’s Mythos Preview and OpenAI’s GPT-5.5 Cyber, scoring 90.9 per cent on the CyberGym benchmark and uncovering more than 200 zero-day holes in open-source code, The Register reported. The trick is teamwork.
The agent pairs Anthropic’s Claude Opus 4.6 with GPT-5.5, and routes each stage of a job to whichever model is best at it. “No single model is best at everything, and none stays state of the art for long,” Wiz’s researchers wrote. It is adding Google’s Gemini Flash Cyber next.
Microsoft’s version, at half the cost
Microsoft is chasing the same idea. Its MDASH harness scored 95.95 per cent on CyberGym, beating Mythos, Gemini and GPT. It splits the work between red-team agents that find exploitable flaws and green-team agents that patch them.
It runs a small in-house model, MAI-Cyber-1-Flash, for up to 90 per cent of tasks, and hands the hardest 10 per cent to the larger GPT-5.4. Pairing the two halves the cost, according to Microsoft AI chief Mustafa Suleyman. The models “deliver better performance than all of the other models combined,” he said, “at 50 per cent of the cost.”
The single-model scores show the gap. On CyberGym, GPT-5.5 Cyber managed 85.6 per cent, and Anthropic’s Mythos 5 hit 83.8 per cent. Google’s Gemini 3.5 Flash Cyber reached 83.2 per cent. Stitched together, the systems beat all of them.
The caveats
A few things temper the hype. The CyberGym figures are the vendors’ own, and Atlas is not for sale, it runs inside Wiz. A benchmark win is not a production track record. Both efforts also sit in a wider wave of bug-hunting AI, so the scores keep moving.
There is a cost problem, too. Pointing a frontier model at a codebase once is expensive, and the result goes stale as code changes by the minute. Wiz’s Nir Ohfeld argues the real test is whether a system keeps finding flaws “continuously and economically” as better models arrive.
CosmosEscape is the proof of concept either way. The same models now finding flaws faster than ever cut both ways. This week Wiz aimed them at Azure and found a master key. The past month is a reminder that attackers can aim them just as easily.
Get the TNW newsletter
Get the most important tech news in your inbox each week.