xAI's Grok Voice Transcribe 2.0 Claims Double the Accuracy at the Same Price

xAI says Grok Voice Transcribe 2.0 is twice as accurate as version 1.0 and tops a 32-model streaming leaderboard, with batch audio at $0.10 per hour and streaming at $0.20 per hour.

Sep 18, 2026
5 min read
Technobezz
xAI's Grok Voice Transcribe 2.0 Claims Double the Accuracy at the Same Price

Don't Miss the Good Stuff

Get tech news that matters delivered weekly. Join 50,000+ readers.

xAI says its new Grok Voice Transcribe 2.0 is twice as accurate as the model it replaces, and it costs the same as Grok Voice Transcribe 1.0. The company also says the model ranks first for accuracy among 32 streaming models on the Artificial Analysis leaderboard. It is built on the audio foundation model behind Grok Voice, the same speech technology used by the Grok assistant in Tesla vehicles.

The pricing is the part most teams will check first. Batch transcription runs $0.10 per hour of audio, while streaming costs $0.20 per hour. On short phrases, xAI reports a word error rate of 6.8%, down from 20.6% for the previous version. The company says the model leads every model tested on telephony, which it sampled at 8 kHz.

Read more: Grok Spreads False Claims About Australia's Bondi Beach Shooting

Multilingual work is where xAI says the biggest gain sits. The model transcribes dozens of languages, detects the language on its own, and follows a mid-recording switch between languages in a single pass. It was trained on live, noisy, multilingual audio from diverse environments and then refined with post-training. xAI says the improvement over version 1.0 holds across four internal production sets.

Existing Speech-to-Text API integrations pick up the accuracy gain without code changes, according to the announcement. The model handles both recorded files and URLs in batch mode, and can transcribe an audio stream in real time. It returns word-level timestamps with start and end times plus confidence scores, and speaker diarization is included at no extra cost.

For teams with more specialized audio, the model can transcribe up to 8 channels independently and accept up to 100 domain terms per request for key term biasing. Its text formatting writes out numbers, dates, currencies, phone numbers, and emails in written form, and it drops filler words such as um and uh. Smart turn detection, which spots the end of a speaker's turn for voice agents, is part of the package.

The announcement ties the model to a workflow involving Atlassian's Loom and Cursor, where Loom transcripts can feed into Cursor for code updates. "We've always believed the best way to move work forward is to capture context once and let it flow everywhere. With Grok powering Loom's speech-to-text and Cursor turning that into code, we're closing the loop from context to code: record what you mean, and the work gets done. It's a glimpse of where AI-assisted development is headed." said Sanchan Saxena, SVP of Teamwork Collection, Atlassian.

Version 2.0 will soon become the default in the Speech-to-Text API, and version 1.0 is set to be deprecated in the coming weeks. Developers who want to stay on the older model during the transition can pin grok-voice-transcribe-1.0. xAI gave no release date for the default switch, and no general availability date or region list.

The release lands two days after xAI detailed memory in Grok Build, which carries conventions, decisions, and project facts across sessions. Earlier in the month the company introduced Grok Bot for enterprises, a persistent-agent product it said found over $100,000 in direct savings on vendor spend.

Share