Microsoft MAI-Transcribe-2: Speed, Accuracy, and Price vs GPT-Transcribe and Scribe v2
Microsoft priced MAI-Transcribe-2 at $0.10 per audio hour, but the fine print reveals two catches procurement teams should read before switching.
AnIntent Editorial
Photo by Amr Taha™ on Unsplash
Most coverage of Microsoft MAI-Transcribe-2 has treated it as a straightforward win: fastest, most accurate, cheapest speech-to-text on the market. That framing comes almost verbatim from Microsoft's own launch post, and it collapses the moment you check the independent leaderboards. On one benchmark Microsoft leads. On another, it finishes second to Alibaba. And the headline $0.10 per audio hour price is not the standard rate at all.
That gap between the marketing and the fine print is the actual story for anyone deciding whether to migrate a transcription pipeline in late 2026.
The Claim Microsoft Made, and the One It Didn't
Microsoft published the launch on September 3, 2026, positioning the release in a single sentence: MAI-Transcribe-2 is billed as the fastest, most accurate, and cheapest speech recognition model in the world, available in public preview through Microsoft Foundry and the MAI Playground. The company says the model ranks first on the FLEURS multilingual benchmark across 60 languages with a 5.2 percent average Word Error Rate, beating Gemini 3.5 Transcribe, GPT-Transcribe, Whisper V3-Large, and ElevenLabs Scribe v2.
What Microsoft did not put in the headline: the model is a preview release. Microsoft's own documentation labels MAI-Transcribe-2 as a public preview with no SLA and explicitly not recommended for production workloads. That caveat is missing from the press messaging, and it changes the calculus for any team weighing a migration from a paid, GA-supported endpoint.
The price tag also comes with a timer. The $0.10 per audio hour figure is a limited-time promotional rate valid through the end of 2026, and Microsoft has not disclosed what the standard post-promotional rate will be. Any total-cost-of-ownership model built on the launch number is a placeholder, not a forecast.
What Actually Counts as Best on the Leaderboards
There are two accuracy numbers in play, and they measure different things. On Microsoft's chosen benchmark, FLEURS across 60 languages, the 5.2 percent WER puts MAI-Transcribe-2 in first place. On the independent Artificial Analysis WER leaderboard tracked by AlphaSignal, the picture shifts. MAI-Transcribe-2 ranks second overall on that leaderboard with a 2.0 percent AA-WER, behind Alibaba's Fun-Realtime-ASR-preview at 1.7 percent AA-WER.
That is not a rounding error. It means the "most accurate in the world" claim only holds if you accept the specific benchmark Microsoft picked. Alibaba's model wins on the neutral scoreboard, and Microsoft's own press page does not mention it.
The speed claim survives closer scrutiny better. Microsoft says MAI-Transcribe-2 processes long-form audio up to 10 times faster than leading competitors, and Artificial Analysis measurements are consistent with that scale. On AA testing MAI-Transcribe-2 hit 410.7 times real-time throughput, second only to Deepgram Nova-3 in raw speed. For a batch transcription workload measured in thousands of hours per day, that speed advantage translates directly to lower compute-time bills upstream of the per-audio-hour rate.
MAI-Transcribe-2 vs GPT-Transcribe on the Numbers
OpenAI has not published a comparable FLEURS score for GPT-Transcribe in the same 60-language configuration Microsoft used, which is part of why Microsoft's chart looks decisive. What the head-to-head actually shows: on Microsoft's benchmark, MAI-Transcribe-2 wins; on Artificial Analysis, GPT-Transcribe does not appear in the top two positions occupied by Alibaba and Microsoft. The MAI-Transcribe-2 vs GPT-Transcribe comparison in September 2026 is a comparison between a preview model priced at a promo rate and a GA model priced at production rates. Those are not the same product category.
The 70 Percent Price Cut That Isn't Really 70 Percent
Microsoft's price story is the most aggressive part of the launch. Neowin's coverage lays out the trajectory: MAI-Transcribe-1 launched in April 2026 at $0.36 per hour covering 25 languages with a 3.9 percent WER on those top languages, MAI-Transcribe-1.5 followed in June expanding to 43 languages with keyword biasing and speed improvements, and MAI-Transcribe-2 now cuts the launch price by more than 70 percent versus the original $0.36 figure.
Against ElevenLabs, the gap is narrower but still real. ElevenLabs Scribe v2 sits at 2.2 percent AA-WER at roughly $0.22 per audio hour, making MAI-Transcribe-2 approximately 54 percent cheaper at the promotional rate. At $0.10 per hour, MAI-Transcribe-2 works out to about $1.67 per 1,000 minutes of audio. For a call-center or media-archive workload processing five million minutes a month, that is roughly $8,350 versus roughly $18,300 at Scribe v2 pricing.
Those savings only hold while the promo does. Independent analysis at DataNorth makes the same point directly: analysts warn procurement teams to treat the $0.10 per hour figure as a pilot price rather than a long-term TCO assumption, given the undisclosed permanent rate. If the standard rate lands anywhere near Scribe v2's $0.22, the switching argument based purely on cost gets much thinner.
The 5.2 Percent Number Hides Where the Model Struggles
Here is the detail buried in the WER math that most launch coverage skipped. Microsoft reports two different WER figures for MAI-Transcribe-2: 5.2 percent averaged across all 60 FLEURS languages, and 3.4 percent for the top 25 languages, meaning aggregate scores can mask weaker performance on lower-resource languages, accents, or specialized domains.
That 1.8-point spread is a signal, not noise. Think of it like a car's combined fuel economy rating: the number is honest but it averages highway cruising with stop-start traffic, and if you drive only one of those modes your real mileage will differ. A logistics company transcribing Swahili customer service calls or an insurer processing rural US telephony audio should not extrapolate from the headline 5.2 percent. The practical recommendation is that organizations with a specific regional or telephony audio mix benchmark their own data before replacing existing pipelines.
This is also the axis where Whisper's ecosystem still matters. Open-weight fine-tunes of Whisper V3-Large on domain-specific corpora, medical dictation, legal deposition, low-resource languages, remain competitive on the workloads that matter most to the buyers of those fine-tunes, even if the base model loses on FLEURS. A locked preview API cannot be fine-tuned. That is a real trade-off, and it is not visible in any leaderboard.
Teams evaluating the best speech-to-text API 2026 options should treat this as the defining question: is your audio close enough to the FLEURS distribution that a general model wins, or specialized enough that a tuned pipeline still wins even at higher per-hour cost?
The Features That Actually Matter for Production Pipelines
Raw WER and price are the marketing story. The feature list is what determines whether a model can drop into an existing pipeline at all. MAI-Transcribe-2 adds speaker diarization, word-level timestamps, configurable clean versus verbatim transcription styles, and keyword biasing over the prior versions.
Those four items map to specific real workloads:
- Speaker diarization is table stakes for meeting transcription, podcast pipelines, and call-center analytics. Its absence in earlier MAI versions kept the model out of those use cases entirely.
- Word-level timestamps are what make caption alignment, video editing snap-to-word workflows, and search-into-audio features work.
- Clean versus verbatim modes matter to legal, medical, and journalism pipelines that need either edited readability or every "um" and "uh" preserved.
- Keyword biasing, which shipped in MAI-Transcribe-1.5, addresses the domain-vocabulary problem, product names, drug names, ticker symbols, that generic models mangle.
GPT-Transcribe, Scribe v2, and Deepgram Nova-3 all offer variants of these. The competitive question is not whether MAI-Transcribe-2 has the features, it is whether Microsoft's implementation of diarization holds up on overlapping speakers in noisy audio, something neither Microsoft's press page nor the current independent benchmarks measure directly.
Where This Fits in Microsoft's Broader AI Push
MAI-Transcribe-2 is a MAI-branded model, part of Microsoft's in-house model family developed independently of OpenAI. Neowin notes the model is also coming to OpenRouter, though it was not yet available there at launch. That distribution move matters more than it looks. Putting a first-party Microsoft model on a neutral aggregator is the same play Google made with Gemini pricing on the API market: buy market share now, monetize the workloads that stick later.
It also reduces Microsoft's dependence on OpenAI as the sole source of frontier-model capability inside Azure. The pattern is visible across the AI infrastructure category, where hyperscalers are quietly building house-brand alternatives to the frontier labs whose APIs they resell. Speech-to-text is a natural first target: the workloads are large, the accuracy plateau is high, and the differentiation is mostly on price and latency rather than capability novelty.
The Microsoft AI transcription model story two years from now will be less about who has the lowest FLEURS score in a given quarter and more about which vendor's SLA and pricing hold steady across a three-year enterprise contract. Neither is settled today.
What to Do With This Information Before End of 2026
If you run a transcription workload and you are looking at MAI-Transcribe-2 pricing accuracy benchmarks to decide whether to switch:
- Do not sign a migration plan around $0.10 per hour. Model your break-even against a plausible standard rate of $0.20 to $0.30 per hour and see if the math still works.
- Benchmark on your own audio, not on FLEURS. The 1.8-point WER gap between Microsoft's 60-language and top-25 numbers is your warning that averages lie.
- Confirm the preview-vs-GA status matters for your compliance posture. No-SLA endpoints are fine for R&D and internal tooling. They are not fine for regulated production.
- Watch for the OpenRouter listing. Access through a neutral aggregator changes the switching cost profile compared to lock-in through Foundry.
The most useful frame for MAI-Transcribe-2 in September 2026 is not "the new leader" but "the aggressive new entrant priced to move." That is a genuinely interesting product. It is not the same thing as the finished, production-safe transcription default that Microsoft's headline implies, and treating it as such is the mistake worth avoiding.
Frequently Asked Questions
Is MAI-Transcribe-2 available for production workloads?
No. Microsoft's documentation labels MAI-Transcribe-2 as a public preview with no SLA and explicitly not recommended for production workloads. Teams needing guaranteed uptime should keep a GA-supported model in the pipeline until Microsoft moves it to general availability.
How does MAI-Transcribe-2 compare to Alibaba's transcription model?
On the independent Artificial Analysis WER leaderboard, Alibaba's Fun-Realtime-ASR-preview leads at 1.7 percent AA-WER while MAI-Transcribe-2 sits second at 2.0 percent. Microsoft's press page does not mention Alibaba, so the "most accurate in the world" claim depends on which benchmark you use.
What was the price of MAI-Transcribe-1 when it launched?
MAI-Transcribe-1 launched in April 2026 at $0.36 per audio hour, covering 25 languages with a 3.9 percent WER on those top languages. MAI-Transcribe-2's $0.10 launch price is more than 70 percent lower than that original rate.
How many languages does MAI-Transcribe-2 cover?
Microsoft rates the model on FLEURS across 60 languages with a 5.2 percent average Word Error Rate. Performance on the top 25 languages is stronger at 3.4 percent WER, indicating weaker accuracy on lower-resource languages in the tail.
When will MAI-Transcribe-2 be available on OpenRouter?
Microsoft has said MAI-Transcribe-2 is coming to OpenRouter but it was not yet available on that platform at the September 3, 2026 launch. No specific listing date has been disclosed by either Microsoft or OpenRouter.
Written by
AnIntent Editorial
AnIntent is an independent technology and automotive publication. Our editorial team researches every article from live primary sources, cross-checks key facts across multiple references, and cites claims inline so readers can verify them directly. We cover smartphones, laptops, EVs, gaming hardware, AI tools, and more — with no sponsored content and no paid placements.