Microsoft will begin ranking artificial intelligence models by safety performance, a move aimed at reassuring customers and helping them navigate an increasingly crowded AI marketplace, the company announced this week.
The new safety metric will be added to Microsoft’s “model leaderboard,” a tool launched earlier this month for users of its Azure Foundry developer platform. The leaderboard currently evaluates models based on quality, cost, and throughput (or response speed), and is already used by tens of thousands of clients. According to Sarah Bird, Microsoft’s head of Responsible AI, the addition of a safety category will allow businesses to “directly shop and understand” which AI models best align with their trust and risk requirements.
Bird told the Financial Times that safety ratings will draw from benchmarks such as Microsoft’s own ToxiGen, which detects implicit hate speech, and the Center for AI Safety’s Weapons of Mass Destruction Proxy benchmark, which evaluates whether an AI model could assist in the creation of harmful agents like biochemical weapons.
The move comes as AI developers and cloud service providers face mounting scrutiny over the deployment of autonomous AI agents and concerns about data misuse, privacy breaches, and potential malicious applications. Microsoft, already a dominant force in the cloud ecosystem alongside Amazon and Google, is positioning itself as a neutral platform in the generative AI space, offering models not only from long-time partner OpenAI but also from competitors like Elon Musk’s xAI and Anthropic.
Last month, Microsoft confirmed it would offer xAI’s Grok models on Azure, even after one Grok variant raised alarm by producing content referencing “white genocide” in South Africa due to an unauthorized code modification. xAI has since revised its monitoring systems, but the incident highlighted the importance of transparency and oversight—issues Microsoft’s new safety metric aims to address.
“The models come in a platform, there is a degree of internal review, and then it’s up to the customer to use benchmarks to figure it out,” Bird noted. She emphasized that there is currently no universal standard for AI safety, though Europe’s upcoming AI Act will require developers to implement safety tests before deployment.
Industry experts say Microsoft’s leaderboard may help companies make more informed choices, but warned against overreliance. “Safety leaderboards can help businesses cut through the noise and narrow down options,” said Cassie Kozyrkov, a former chief decision scientist at Google. “The real challenge is understanding the trade-offs: higher performance at what cost? Lower cost at what risk?”
Microsoft has also developed tools to stress test AI systems before they’re deployed. In April, it launched an “AI red teaming agent” capable of automatically simulating attacks to identify vulnerabilities based on user-defined risk criteria.
Despite Microsoft’s effort to champion responsible AI practices, some providers—including OpenAI, as previously reported by the Financial Times—have scaled back on safety investments, citing efficiencies that allegedly don’t compromise security. Bird declined to comment specifically on OpenAI but stressed that developing high-quality AI models requires significant investment in evaluation and risk mitigation.
As regulatory oversight intensifies and AI adoption accelerates, Microsoft’s decision to elevate safety to the forefront of its model rankings may prove pivotal—not just for its own reputation, but for setting a broader industry precedent. Still, experts caution that metrics alone cannot substitute for comprehensive governance.
“Safety metrics are a starting point, not a green light,” said Kozyrkov. “The real work begins after the leaderboard.”

