>Grok STT promises a ten-cent audio hour
ALTIOR AI ADVANTAGEWhat to remember
Grok STT on OpenRouter

The ten-cent transcript test

OpenRouter says an audio hour now costs $0.10 to transcribe, with 25-language coverage, word-level timestamps and speaker diarization.

An hour-long waveform resolving into timed words and separated speakers around a ten-cent price motif.

Ten cents for an hour of transcription is an arresting published price. Yet cheap text is only useful when we can trust the words, timing and speaker labels that come back.

OpenRouter’s launch post establishes availability, features and price. It does not establish how Grok STT performs across our voices, accents, languages or difficult recordings.

Why it matters

Cheap text is not the whole story

The structure around a transcript can matter as much as the words inside it.

Raw audio becomes timed words and separated speakers, enabling search, captions, quotes and edits.
Conceptual flow: audio becomes timed, speaker-separated text that can support search, captions, quotation and editing. This illustrates the stated features, not measured performance.

A plain transcript gives us words. Word-level timestamps can tell us when each word occurred, while speaker diarization can separate who said what.

That structure can make a long recording easier to search, caption, quote and edit. OpenRouter states that Grok STT provides both features; the announcement does not demonstrate how reliably they work.

The launch claim

What OpenRouter actually says

Five concrete details are published; the performance questions remain open.

Five OpenRouter launch claims arranged around a central audio waveform.
OpenRouter says Grok STT is live with speech-to-text in 25 languages, word-level timestamps, speaker diarization and a published price of $0.10 per audio hour.
A verification prism routes an audio transcript toward four unresolved performance and availability questions.
The launch post provides no benchmark for accuracy, latency, noisy-room handling, per-language quality or diarization quality.
$0.10per audio hour
25stated languages
Word-leveltimestamps stated
Speakerdiarization stated

OpenRouter says Grok STT is now live on its platform. The same post states speech-to-text in 25 languages, word-level timestamps, speaker diarization and a price of $0.10 per audio hour.

Those are specific, useful claims, but they are not a quality assessment. We do not yet have evidence here for transcription accuracy, response time, noisy-room performance, consistency across languages or the quality of speaker separation.

Source proof

The launch receipt

The post and its attached visual establish provenance, with one naming caveat.

Provider image: source-001-image-01.jpg
OpenRouter’s attached launch visual uses the label “Grok Voice API”, while the post itself announces “Grok STT”. The source does not establish that these labels are exact synonyms.

The attached visual is useful proof that the announcement came from OpenRouter, but its label needs care. The post names Grok STT; the graphic says “Grok Voice API”.

We can preserve both facts without collapsing them into one product identity. The written announcement remains the source for the Grok STT availability, feature and price claims.

Grok STT from @SpaceXAI is now live on OpenRouter. Speech-to-text in 25 languages with word-level timestamps and speaker diarization. $0.10 per audio hour.

OpenRouter on X
Evidence boundary

Proof, with limits

A concise announcement can prove what was said without proving how well the service performs.

The post is direct enough to support four points: availability on OpenRouter, a stated 25-language scope, word-level timestamps and speaker diarization, plus the published hourly price.

It contains no test result, comparison or real-world transcript. Accuracy, latency, noise handling, language-by-language quality and speaker-separation quality therefore remain unanswered.

How it works

How structured transcription works

The promise is a transcript that preserves both timing and speaker structure.

Conceptual flow from audio input through recognition, timed words and speaker separation to searchable, caption-ready output.
Conceptual sequence only: audio enters a speech-to-text system, recognised words receive timings, speakers are separated, and the resulting transcript can support search, captions and edits.

Speech recognition turns an audio signal into words. Word-level timestamps attach timing information to those words, while diarization groups the speech by speaker.

Together, those layers can give us something more useful than a continuous text block. This is a conceptual explanation of the features OpenRouter names, not a capture of Grok STT’s output.

Recognise

Structure

Use

Practical consequence

What this could change for us

A low published price could make more recordings worth processing—if the output survives our own tests.

Creator workflow from upload through structure, search, captions and evaluation, with quality, billing and availability checks.
A neutral workflow from audio upload to structured transcript, search, captioning and editing. The downstream uses are practical possibilities, not demonstrated Grok STT results.

At $0.10 per audio hour, the published price could make transcription economical for interviews, meetings, podcasts and larger archives. Timestamps and speaker separation could then reduce the work needed to find a quote, cut a clip or prepare captions.

But the useful cost is not the API price alone. If we spend time correcting words, rebuilding speaker labels or checking difficult passages, the workflow changes. That is why our own audio—not the launch claim—has to decide whether the service fits.

Upload

Inspect

Decide

The decision

The real test is your audio

Price and features earn a trial; representative recordings determine whether the output is useful.

The launch case is narrow but credible: OpenRouter has published a striking price alongside a useful transcription feature set. What we still need is evidence from the conditions that matter to us.

A fair test should include our usual voices, accents, languages, overlaps and background noise. We can then inspect the words, timestamps and speaker assignments—and count the correction work—before treating the ten-cent hour as a workflow advantage.

The ten-cent price makes Grok STT easy to test. Only our recordings can show whether it is cheap to use.

Altior analysis of the OpenRouter announcement

Stress-test Grok STT on real audio

Transcribe the attached audio verbatim. Preserve the spoken language, separate each speaker consistently, and include a timestamp for every word. Do not correct grammar, remove filler words, infer unheard speech, translate the recording, or merge uncertain speakers. Mark any uncertain word as [unclear] with its timestamp. Return: (1) the complete speaker-labelled transcript; (2) a table of every [unclear] segment; and (3) the total audio duration and number of distinct speakers detected.
Ready to copy
ALTIOR AI ADVANTAGE
What comes next

Watch the evidence arrive

Track independent tests and OpenRouter’s product details, then run representative audio before making Grok STT part of a serious workflow.

Try the prompt

Six questions still open

  • Accuracy tests
  • Timing and latency
  • Difficult audio
  • Language and billing detail