Microsoft AI launched MAI-Transcribe-2 on 3 Sep 2026 as a public preview, priced at $0.10 per hour of audio until 31 Dec 2026. It's a speech-to-text model with speaker labels and word-level timestamps built in. If you transcribe interviews, podcasts, or client calls, that's a cheap window to test it. It's also a short one, and Microsoft hasn't said what the price will be in January.
For MAI-Transcribe-2 freelancers, the call is simple. Try it if you transcribe every week, want speaker labels and timestamps you can quote from, and don't mind calling an API or using a developer playground. Skip it if your meeting tool or editor already gives you transcripts that are good enough. Wait if you need settled production terms, a known 2027 price, or a no-code app. The Streaming version is only for live captions, and it drops speaker labels and timestamps.
Affiliate disclosure: some tool or platform links may later be affiliate or referral links. Product facts below come from Microsoft AI's MAI-Transcribe-2 model page, its 3 Sep launch post, and Microsoft's Foundry blog on Tech Community, as of 6 Oct 2026 (SAST). It's a preview, so prices, features, and availability can change. Re-check before you put client audio through it. Not legal advice. Not a guarantee of clients, savings, or income.
What MAI-Transcribe-2 is
MAI-Transcribe-2 is a developer model, not an app. You send it audio and it sends back text. Microsoft says it's available in public preview in Microsoft Foundry and through Azure Speech, and its launch post says you can demo it through Foundry, the MAI Playground, and OpenRouter.
There's no app to download and no Record button. If you've only ever pressed "transcribe" inside Zoom, Meet, or an editor, this is one layer down.
It follows MAI-Transcribe-1.5, which the model page still lists. The older model is listed at $0.36 per audio hour, supports 43 languages, and has no speaker labels or word timestamps. That gap is most of the reason a freelancer would care about version 2.
The price, and what nobody knows yet
Here's what Microsoft publishes:
- MAI-Transcribe-2: $0.10 per hour of audio, introductory. Microsoft's Foundry post says that price runs through 31 December 2026. The launch post calls it "a limited-time offer until the end of the year".
- MAI-Transcribe-2-Streaming: $0.54 per hour of audio, also marked introductory on the model page. No end date is given.
- MAI-Transcribe-1.5: $0.36 per hour of audio, for comparison.
Here's what Microsoft doesn't publish, so we won't guess:
- The price after 31 Dec 2026. It may stay, rise, or change shape. Nobody outside Microsoft knows.
- When the $0.54 streaming price ends.
- Free tier or trial credits. Not stated on the pages we checked.
- File-size or length limits.
- OpenRouter's own per-hour price. Check OpenRouter's model listing yourself if you go that route.
If you work in rands, convert at the day's rate yourself. If you already budget API usage, the same habit applies here. Our API cost guide for freelancers covers the mindset: log a real job, read the meter, then decide.
Features freelancers will actually use
These are the features that matter if you bill for transcripts, quotes, or show notes:
- Speaker diarization. The model separates speakers and attributes words to the right person. For a two-person interview, that's the difference between a usable transcript and an hour of relabelling.
- Word-level timestamps. Every word gets a time. That helps when you need to find the quote at 34:12 or cut a podcast clip.
- Clean or verbatim style. "Verbatim" keeps fillers and false starts. "Clean" removes them for readable notes and published transcripts. Pick verbatim for research and anything legal-adjacent, clean for show notes.
- Keyword biasing. You can feed it terms it should expect: client names, product names, jargon, abbreviations.
- 60 languages, with automatic language identification and code switching. Microsoft's examples are Hinglish and Spanglish. Microsoft also says it holds up in noisy audio.
One caution on languages. Microsoft's pages don't say anything specific about South African English accents, isiZulu, or Afrikaans. Don't assume either way. Test on your own audio.
How you'd actually use it
Microsoft names four routes: Microsoft Foundry, Azure Speech, the MAI Playground, and OpenRouter.
Microsoft's launch pages don't give a step-by-step setup for freelancers, so we won't invent one. In practice the playground and OpenRouter routes look like the lightest way to test a file, and Foundry or Azure Speech make sense if you already have an Azure account. Either way you'll be dealing with developer screens, keys, and usage meters, not a consumer app.
Two things the pages don't answer, and you should check before uploading anything real:
- Which Azure regions serve it. Microsoft doesn't list regions on the launch pages. We're not saying whether it's offered in South Africa or anywhere else. Check inside Foundry or Azure for your own account.
- Data handling and preview terms. Microsoft's launch pages don't spell out retention or privacy terms for your audio, or production terms during the preview. Read the Azure preview terms and the data terms for the route you pick. If a client contract covers recordings, that comes first.
Microsoft's accuracy and speed claims
These are Microsoft's numbers, mostly citing evaluations by Artificial Analysis. We haven't tested them.
- First on FLEURS across 60 languages, with an average word error rate of 5.2%, per the launch post.
- Second on the Artificial Analysis word-error-rate leaderboard, and leading the Artificial Analysis accuracy-latency frontier.
- Speed: the model page lists an hour of audio transcribed in about 10 seconds end to end. The launch post says it's 10x faster than OpenAI's GPT-Transcribe, 7x faster than ElevenLabs' Scribe v2, and 5x faster than Gemini 3.5 Transcribe.
- The Whisper question: Microsoft's launch post names Whisper V3-Large among the models it says MAI-Transcribe-2 beats. That's Microsoft's comparison, not ours.
One oddity: the model page's accuracy blurb mentions "MAI-Transcribe-2.1". The launch posts say MAI-Transcribe-2, so we do too.
Benchmarks are averages. A 5.2% average says nothing certain about a noisy Zoom call with three people talking over each other.
If you already use ElevenLabs for voice work, our ElevenLabs for freelancers guide has the context. The Scribe speed comparison above is Microsoft's claim, not ours.
Try / Skip / Wait sheet
Try if you transcribe regularly (interviews, podcasts, client calls), you need speaker labels and word timestamps, and you can use Foundry, the MAI Playground, OpenRouter, or an API. Test before 31 Dec while the $0.10 launch price runs.
Skip if your meeting tool or editor already gives you transcripts that are good enough, and you don't want any API setup. If that's Descript, our Descript vs CapCut comparison covers the editor side.
Wait if you need confirmed production terms, a known 2027 price, or a no-code app. It's a public preview developer model.
Streaming ($0.54/hr) only if you need live captions. The model page says it has no speaker labels, no word timestamps, no keyword biasing, and no style setting.
| Your situation | Try | Skip | Wait |
|---|---|---|---|
| Weekly interviews or podcasts, comfortable with an API or playground | Yes: one non-client file before 31 Dec | If your current transcripts already need no fixing | — |
| Meeting tool or editor transcripts are good enough | Only out of curiosity | Yes: no setup needed | — |
| You want a no-code app | — | For now | Yes: it's a developer model |
| Confidential client recordings | Only after reading the data and preview terms | If the contract rules out third-party processing | Yes, until terms are clear to you |
| Building a 2027 transcription budget | Test now, budget later | — | Yes: post-December price not published |
| Live captions for a webinar or stream | Streaming, if speaker labels don't matter | If you need labels or timestamps live | — |
Get the 1-page Try / Skip / Wait transcription chooser
Optional printable of this free guide. Soft link until checkout. Not a "paid summary." Not a guarantee of clients, savings, or income.
The Joburg reality check
Pick a recording you own and don't mind sharing, like a podcast episode you've already published. Not a client call. Run it through the model and through whatever you use now, then compare speaker labels, names, and the quotes you'd actually use.
Upload over fibre you trust, not mobile data halfway through load-shedding. And before any client audio goes anywhere, read the data terms for the route you chose and check your contract. That's how we'd test it from a Joburg desk. It works the same anywhere.
FAQ
Is MAI-Transcribe-2 free?
Not that we can see. Microsoft lists an introductory price of $0.10 per hour of audio. Its pages don't mention a free tier or trial credits.
What happens to the price after December?
Unknown. Microsoft says the $0.10 rate runs through 31 Dec 2026 and hasn't published what comes next.
Does the Streaming version have speaker labels?
No. Microsoft's model page lists MAI-Transcribe-2-Streaming without diarization, word-level timestamps, keyword biasing, or transcription styles.
Can I use it without coding?
Partly. You can demo it in the MAI Playground, and it's listed on OpenRouter. For regular work you'll still be using developer tools, not a consumer app.
Is it available in South Africa?
We don't know. Microsoft's launch pages don't list Azure regions. Check in Foundry or Azure for your own account before you plan around it.
Sources: MAI-Transcribe-2 model page (Microsoft AI), MAI-Transcribe-2 is the fastest, most accurate and cheapest speech recognition model in the world (Microsoft AI, 3 Sep 2026), MAI-Transcribe-2: Highest quality transcription, at the fastest speed and lowest cost (Microsoft Foundry Blog, Tech Community, 3 Sep 2026). Re-check price, end date, and availability at paste time.
