
Microsoft AI On Wednesday, it unveiled two new in-house models – MAY-Image-2.5-Prothe highest resolution image generator to date and MAI-Voice-2-Flashspeech model built for high-volume enterprise workloads — while publishing production data, the company’s most aggressive argument that OpenAI can power its products without relying on boundary models.
announcement made by Microsoft AI’s Superintelligence teamIt comes nearly a year after the company committed to building purpose-built models internally, and it reaches an unusual level of specificity about where those models now work: Bing, PowerPoint, OneDrive, Dynamics 365, Excel, GitHub Copilotand Azure. The message to enterprise buyers — and by extension OpenAI — is that Microsoft’s native models are no longer research projects. They are production infrastructure serving millions of users.
"Each of these improvements is a step toward the same goal: Microsoft products equipped with Microsoft models," the company wrote on its announcement blog.
How MAI-Image-2.5-Pro and MAI-Voice-2-Flash separate opposite sides of the AI value curve
The two new releases occupy opposite ends of what Microsoft calls the quality-speed-cost curve, and the placement is intentional. MAY-Image-2.5-Pro aims for premium: hero images, detailed editing, and accurate image text rendering – the latter of which has long been a notorious weak point for image creation models. Microsoft priced this model at $5 per million text input tokens, $8 per million image input tokens, and $106 per million image output. Base MAI-Figure-2.5 Recently released model #2 for image editing in Arenaa community leaderboard that has become a virtual leaderboard for generative media.
The creative industry is gaining attention. Advertising giant WPP’s global chief creative officer, Rob Reilly, called the Pro model "Strong breakthrough for GenMedia tools" It was added in a statement included in Microsoft’s announcement "Microsoft has firmly established itself among the leaders in the field of generative artificial intelligence."
MAI-Voice-2-Flash it goes in another direction. First viewed on Microsoft Set up a conferenceFlash is twice as fast as MAI-Voice-2 and costs 32% less, at $15 per million characters. It is designed for the unusual but huge high-volume voice market – call centers, voice agents and real-time speech applications – where latency and cost per call are more important than marginal gains in expressiveness. Together, the two models reflect a strategy of building model families rather than a single flagship, because, as the company says, a creative studio pursuing maximum fidelity has very different needs than a customer service operation handling millions of calls a day.
Microsoft’s manufacturing benchmarks show internal models that reduce GPU costs by up to 89%
The presentation of the model is less newsworthy than the placement dimensions that Microsoft has attached to them – numbers that read like a systematic case for changing third-party boundary models in the product portfolio.
Bing Image Creator it’s totally working now MAI-Figure-2.5marks the first time a consumer imaging tool has been completely built-in, end-to-end. In PowerPoint, Microsoft says that MAI-Image-2.5 reduces GPU costs by up to 84% compared to OpenAI’s image model, GPT-Image-2. On OneDrive, where MAI-Image-2.5 is now the default for basic image editing scenarios, the company reports a 26% increase in save rates, about 25% lower P95 latency, and 2.5x greater efficiency in medium-use production workloads.
On the sound side, MAI-Voice-2-Flash is now deploying Dynamics 365 Contact Center – a platform used by customers including T-Mobile and EasyJet – where Microsoft claims GPU costs have been reduced by up to 89%. The model is also integrated into Azure Voice Live for developers building speech-to-speech agents.
Perhaps the most consequential placement sits in healthcare. Microsoft’s Dragon CopilotUsed by 170,000 healthcare providers and responsible for processing 28 million patient encounters last quarter, it is now working on MAI-Transcribe-1.5 for multilingual workflows in 58 languages. Microsoft says that internal evaluations show a 50% relative reduction in both transcription and language identification error rates in most languages—a meaningful claim in a domain where transcription errors can spread directly to clinical records.
Inside the “hill-climbing” strategy that allows small models to defeat GPT-5.6 in Excel
In a follow-up post published the same day, Microsoft detailed the methodology behind these results — which it called "mountain climbing car," an integrated flywheel of data, models and product "trailer" that covers them.
It is the most obvious example MAI-Code-1-Flashthe lightweight coding model was launched on GitHub Copilot in June. Microsoft says the model achieves about 10% higher code acceptance rates than VS Code’s GPT-5.4 Mini and Claude Haiku 4.5, while using 10% fewer median tokens. Developer retention tells a similar story: users were 6% more likely to return within days compared to GPT-5.4 Mini, and 11% more likely to return to Claude Haiku 4.5.
Then Microsoft did something even more interesting. MAI-Code-1-Flash went through the checkpoint and beyond trained him in an Excel reinforcement learning environmentteaching the coding model the tools and workflows of spreadsheet knowledge work. According to production user feedback, the result is a model on par with GPT-5.6 for the most common Excel tasks – while also being small enough to run on Nvidia’s older H100 and even A100 GPUs instead of requiring the latest generation of accelerators.
This hardware detail is worth highlighting. Every major AI company is vying for cutting-edge chip allocation, and a model that delivers borderline quality on two generations of old silicon is fundamentally changing the economics of deployment. It also frees up the newest hardware, including Microsoft’s current GB200 cluster, for training instead of servicing.
Satya Nadella’s ‘frontier diffusion’ manifesto reimagines the OpenAI relationship
Microsoft CEO Satya Nadella made the announcements in a lengthy post titled X "Boundary Diffusion and Control," it functions as something close to a strategic manifesto. "We can now take saturated frontier capabilities and deliver them at scale and at lower cost through models optimized for high-use products, while continuing to use frontier models for frontier needs." Nadella wrote, adding that Microsoft "we start redirecting traffic from our first-party surfaces to MAI when our models match or outperform boundary alternatives."
Translated from executive prose: capabilities that were state-of-the-art a year ago are now desktop assets, and Microsoft believes it can replicate them cheaply for the specific, repetitive tasks that dominate real product use. Why pay border prices for a border model when the user just wants to reformat a table column?
Nadella was careful to point this out "Boundary models from OpenAI and Anthropic are part of the orchestration system along with MAI" — but it also expresses the exact principle of model independence, of the company’s valuations "it must continue to climb the hill even if any model is removed."
“What gives Microsoft control is keeping attachment, memory, context and skills out of the model. The subtext is hard to miss.” Reuters reported in April that Microsoft Exclusive license to OpenAI technology the non-exclusive agreement had been renegotiated, and The Information reported last September that Microsoft had It began to incorporate anthropic models included in some products. Wednesday’s announcement completes the triangle: Microsoft as the orchestrator, with its partners’ frontier models as interchangeable components and its own models absorbing an ever-larger share of daily traffic.
Developers welcome cheaper task-specific models, while skeptics question Microsoft’s track record.
The online response captured both appeal and skepticism about the strategy. "I love that people use small models for niche tasks," an X user wrote, @mavihskHe responded to Nadella’s post. "Why do I need to use an omniscient model to change my field in Excel?" Another user, @nabu_linesneatly distilled the pitch: "cost and performance improve when you stop overusing the largest model."
Others were less charitable about Microsoft’s execution record. "Microsoft is the worst at listening to user feedback," the designer wrote @designedbyabinthe company argues "many times they will lose the AI race because they don’t understand user needs." And a user, @tokenoverflowoffered a more dry critique of the model-independence pitch: "after removing microsoft i want it to go up hill."
Skeptics raise a fair point. Microsoft’s self-reported metrics — receive rates, save rates, GPU save — come from its own internal evaluations, not independent benchmarks, and which comparisons the company chooses to publish.
But the logic of the strategy does not depend on any single number. Nadella develops this software "real marginal cost for the first time" Explains why Microsoft is obsessed with tokens, GPUs, and service costs: An 84% reduction in GPU costs is not optimization when AI features run on every keystroke across a billion users’ product portfolio. It’s the difference between a reliable business and a money pit.
Why is Microsoft turning its internal AI playbook into an Azure product?
The last part of the strategy is that Microsoft sells the playbook, not just the models. Nadella clearly laid out his hill-climbing approach "template for every other AI native, SaaS or Enterprise company," and packages Microsoft Foundry and its toolchain, which it calls Frontier Tuning—allowing enterprises to build specialized models against their proprietary assessments and reinforcement learning environments. It turns Microsoft’s internal cost-cutting exercise into an Azure product, giving enterprise customers a reason to run AI workloads in Microsoft’s cloud, even if the models themselves come from elsewhere.
Company emphasis on trained models "in pure, traceable, enterprise-grade data without distillation from third-party models" serves the same commercial purpose. In an industry facing increasing scrutiny over the origins of training data, Microsoft is betting that enterprise buyers and courts will pay attention to where its model capabilities come from. Microsoft said it is now expanding its hill-climbing approach Copilot Chat, Outlookand PowerPointand a public preview is available through both new models Microsoft Foundry and MAI Playground. "None of this is the end point," the company wrote. "We are just getting started."
seven years ago, Microsoft bet more than 13 billion dollars OpenAI will build the future of artificial intelligence. Wednesday’s announcement suggests the company has learned a less expensive lesson since then: The future of AI may belong to whoever builds the frontier, but the profits belong to whoever makes it commonplace.




