Grok's new transcriber is No. 1 on accuracy. It waits half a second
2026-09-19 · 4 min read
On Thursday xAI, which now posts as @SpaceXAI, introduced Grok Voice Transcribe 2.0 and called it "the world's most accurate speech transcription model." The thread had passed 2.8 million views by Friday. Four hours after it went up, @ArtificialAnlys, the independent benchmark shop, had run the model and confirmed the headline for live audio: first place among 32 streaming models. The same post recorded the trade-off, and that is the part worth reading if you run a phone line.
What xAI shipped
The pitch in the thread was plain. The new model is "twice as accurate as its predecessor in customer-support calls, spoken credentials, and short voice commands," per @SpaceXAI, and the price did not move: 10 cents per hour of recorded audio, 20 cents per hour for live streaming. Speaker labels, word timestamps, and custom vocabulary come with it instead of being sold as add-ons.
Among the follow-ups, the one about Atlassian drew the most views. @SpaceXAI said the company tested the model inside Loom, "found it more accurate at capturing user instructions," and now uses it to transcribe every Loom video. The demo shows someone dictating a change request into Loom and exporting the transcript straight into the Cursor coding tool.
A few details in xAI's own write-up matter more than the headline. The company tested on four sets of real audio: 8 kHz customer-support calls, conversations with Grok, people reading phone numbers, email addresses, and street addresses aloud, and short voice commands in 19 languages. The biggest jump came on the multilingual commands, where the word error rate fell from 20.6 percent to 6.8 percent. xAI says it leads every model it tested on the phone-call set, though that is the vendor grading its own homework. Version 1.0 is being retired, and the docs let you pin the old model if you need time.
First on the chart, and the slow one at the top
Artificial Analysis measured a 2.7 percent word error rate on the final streaming transcript, "at 0.49s after end of speech," and gave the model first place on both the final transcript and the first partial one. On recorded audio it scored 2.3 percent, fifth of 59, behind Microsoft's MAI-Transcribe-2 and ElevenLabs Scribe v2.
The number I would circle is the half second. The two models just behind it on streaming accuracy, Meta's Muse Voice Transcribe and ElevenLabs Scribe v2 Realtime, hand back their text in about 0.15 seconds. Grok Voice Transcribe 2.0 is more accurate and roughly three times slower to return it, and it is slower than its own predecessor too. xAI bought accuracy with latency.
My read
For recorded audio the trade-off does not exist. Nobody cares whether a voicemail transcript arrives in a tenth of a second or half a second, and at a dime an hour, transcribing every call a small office took this year costs about what lunch does. The model at the top of the chart is also one of the cheapest on it, which does not happen often.
For live voice, the half second is real, because a caller hears it as a pause. A voice agent that waits for the final transcript before deciding what to say will feel a beat slower than one built on a faster, slightly sloppier model, and callers judge a phone bot by that beat more than by a missed word. So the pick comes down to which mistake costs you more: mishearing a callback number, or sounding like a machine thinking.
The credentials set deserves a pin too. xAI built a whole test out of people reading account numbers and addresses aloud, because that is what support audio is full of. It is also the audio most businesses would least want sitting on someone else's servers. If you send calls to any transcription API, the data retention terms deserve a closer read than the benchmark chart.
What this means for your business
At these prices the transcript itself is close to free. The work is in deciding which calls to capture, what the transcript should turn into, and who reads it. A few places it pays off quickly:
- Intake calls that become a draft job ticket with the address and callback number already filled in, checked by a person before it goes out
- After-hours voicemails summarized into one morning list instead of a queue to replay
- A custom term list, up to 100 per request, loaded with your product names, technician names, and the streets you serve
Illinois requires everyone on a call to consent to recording, so that notice goes in before any of this. And pilot the model on twenty of your own calls before you trust a leaderboard. The 8 kHz support calls in xAI's test set are not the same as your front desk on a Tuesday.
If you want help deciding which of your calls are worth capturing first, and what the transcript should do once it exists, that is the kind of question New Face Design's free process audit is built for. Start here.