The ten-cent transcript test
OpenRouter says an audio hour now costs $0.10 to transcribe, with 25-language coverage, word-level timestamps and speaker diarization.

Ten cents for an hour of transcription is an arresting published price. Yet cheap text is only useful when we can trust the words, timing and speaker labels that come back.
OpenRouter’s launch post establishes availability, features and price. It does not establish how Grok STT performs across our voices, accents, languages or difficult recordings.
Cheap text is not the whole story
The structure around a transcript can matter as much as the words inside it.

A plain transcript gives us words. Word-level timestamps can tell us when each word occurred, while speaker diarization can separate who said what.
That structure can make a long recording easier to search, caption, quote and edit. OpenRouter states that Grok STT provides both features; the announcement does not demonstrate how reliably they work.
What OpenRouter actually says
Five concrete details are published; the performance questions remain open.


OpenRouter says Grok STT is now live on its platform. The same post states speech-to-text in 25 languages, word-level timestamps, speaker diarization and a price of $0.10 per audio hour.
Those are specific, useful claims, but they are not a quality assessment. We do not yet have evidence here for transcription accuracy, response time, noisy-room performance, consistency across languages or the quality of speaker separation.
The launch receipt
The post and its attached visual establish provenance, with one naming caveat.

The attached visual is useful proof that the announcement came from OpenRouter, but its label needs care. The post names Grok STT; the graphic says “Grok Voice API”.
We can preserve both facts without collapsing them into one product identity. The written announcement remains the source for the Grok STT availability, feature and price claims.
Grok STT from @SpaceXAI is now live on OpenRouter. Speech-to-text in 25 languages with word-level timestamps and speaker diarization. $0.10 per audio hour.
OpenRouter on X
Proof, with limits
A concise announcement can prove what was said without proving how well the service performs.
The post is direct enough to support four points: availability on OpenRouter, a stated 25-language scope, word-level timestamps and speaker diarization, plus the published hourly price.
It contains no test result, comparison or real-world transcript. Accuracy, latency, noise handling, language-by-language quality and speaker-separation quality therefore remain unanswered.
How structured transcription works
The promise is a transcript that preserves both timing and speaker structure.

Speech recognition turns an audio signal into words. Word-level timestamps attach timing information to those words, while diarization groups the speech by speaker.
Together, those layers can give us something more useful than a continuous text block. This is a conceptual explanation of the features OpenRouter names, not a capture of Grok STT’s output.
Recognise
Structure
Use
What this could change for us
A low published price could make more recordings worth processing—if the output survives our own tests.

At $0.10 per audio hour, the published price could make transcription economical for interviews, meetings, podcasts and larger archives. Timestamps and speaker separation could then reduce the work needed to find a quote, cut a clip or prepare captions.
But the useful cost is not the API price alone. If we spend time correcting words, rebuilding speaker labels or checking difficult passages, the workflow changes. That is why our own audio—not the launch claim—has to decide whether the service fits.
Upload
Inspect
Decide
The real test is your audio
Price and features earn a trial; representative recordings determine whether the output is useful.
The launch case is narrow but credible: OpenRouter has published a striking price alongside a useful transcription feature set. What we still need is evidence from the conditions that matter to us.
A fair test should include our usual voices, accents, languages, overlaps and background noise. We can then inspect the words, timestamps and speaker assignments—and count the correction work—before treating the ten-cent hour as a workflow advantage.
The ten-cent price makes Grok STT easy to test. Only our recordings can show whether it is cheap to use.
Altior analysis of the OpenRouter announcement
Stress-test Grok STT on real audio
Transcribe the attached audio verbatim. Preserve the spoken language, separate each speaker consistently, and include a timestamp for every word. Do not correct grammar, remove filler words, infer unheard speech, translate the recording, or merge uncertain speakers. Mark any uncertain word as [unclear] with its timestamp. Return: (1) the complete speaker-labelled transcript; (2) a table of every [unclear] segment; and (3) the total audio duration and number of distinct speakers detected.Ready to copy
Watch the evidence arrive
Track independent tests and OpenRouter’s product details, then run representative audio before making Grok STT part of a serious workflow.
Try the promptSix questions still open
- Accuracy tests
- Timing and latency
- Difficult audio
- Language and billing detail