IBM releases Granite Speech 5.0 with record transcription speed
New CapabilitiesThe 470M-parameter encoder-only models drop translation and keyword biasing for roughly 20x faster English transcription.
Yesterday: IBM releases Granite Speech 5.0 TurboCTCNew here? Follow stories to track developments over time. Create a free account to get updates when stories you care about change.
Overview
Updated 45 minutes agoIBM says two new models turn spoken English into text more than twelve thousand times faster than real time — about 3.5 hours of audio per second on one Nvidia H200 GPU. Each compact Granite Speech 5.0 model packs just 470 million parameters.
The speed comes from a blunt design choice. IBM removed the language-model decoder that gave earlier Granite Speech releases translation and keyword biasing, leaving an encoder-only pipeline that is over 20 times faster than before. The headline figures are vendor-reported and await independent leaderboard confirmation.
Why it matters
Real-time transcription on a laptop or phone, with no cloud round-trip and no per-minute fees — if the claimed speed survives independent testing.
Questions about this story
Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.
No questions yet — be the first to ask.
Key Indicators
Voices
Curated perspectives — historical figures and your fellow readers.
Play
Exploring all sides of a story is often best achieved with Play.
Higher or Lower
A number from this story, against one from elsewhere in the news — guess which is bigger, then keep the chain going. 5 rounds, 3 strikes; a miss costs a strike and resets your streak.
Keyboard: ↓/L lower · ↑/H higher
0 points — sign up to put that on the leaderboard.
Connections
Sixteen names from the news. Find the four hidden groups of four. Four mistakes max.
Sign up to keep a daily streak — a new puzzle lands every day.
Exit debate?
Your progress in this debate will be lost.
- 1 Two AI personas square off on this story.
- 2 You predict who'll win each round — correct picks earn XP.
- 3 One crossfire question is yours to fire. Pick it carefully.
Couldn't generate a topic
Select Your Champions
Choose one persona for each side of the debate
DEBATE TOPIC
Choose personas with different perspectives for a more dynamic debate.
Select debater for this side:
No debate personas available right now.
Select debater for this side:
No debate personas available right now.
Who's Got This Round?
Make your prediction before the referee scores
The referee scores both sides on
Round Results
Set the Crossfire
Pick the question both personas must answer in the final round
Debate Oracle! You called every round!
Sharp Instincts! You know your debaters!
The Coin Flip Strategist! Perfectly balanced!
The Contrarian! Bold predictions!
Inverse Genius! Try betting the opposite next time!
XP Breakdown
Prediction History
People Involved
Organizations Involved
IBM's Granite Speech line produces compact ASR and speech translation models; the 5.0 models are built for speed on edge devices.
NVIDIA's H200 data-center GPU is the reference hardware for IBM's throughput claims.
Timeline
-
IBM releases Granite Speech 5.0 TurboCTC
Latest Product releaseIBM launched two 470M-parameter English ASR models claiming 12,600 RTFx on an H200. Official OpenASR Leaderboard rankings were still pending at publication.
Historical Context
3 moments from history that rhyme with this story — and how they unfolded.
Mozilla DeepSpeech (2017)
Mozilla shipped DeepSpeech, an open-source, connectionist-temporal-classification (CTC) speech recognizer designed to run on phones and low-power hardware without a cloud connection.
Became a reference for on-device, self-contained speech recognition.
Demonstrated that encoder-only CTC models could deliver practical, offline transcription.
IBM's encoder-only design revives that same philosophy with token-level output and claims of far higher throughput.
Distil-Whisper (2023)
Hugging Face distilled OpenAI's Whisper large-v2 into a 756-million-parameter model that ran 6x faster and was 51% smaller while staying within 1% word error rate on most benchmarks.
Showed that distillation could shrink a massive speech model while keeping accuracy nearly intact.
Set the expectation that speed and size gains did not have to cost much accuracy.
IBM's 5.0 release pushes the same speed-versus-accuracy trade-off far harder — a 20x throughput jump instead of 6x.
NVIDIA Parakeet (2024)
NVIDIA released Parakeet TDT models through its OpenASR toolkit — non-autoregressive, edge-focused speech recognition models that climbed to the top of open ASR leaderboards.
Gave developers a fast, compact open alternative to Whisper-class models.
Made the OpenASR and FFASR leaderboards the de facto benchmark arena for compact ASR.
Granite Speech 5.0 enters that same arena, ranking among the fastest two models on the FFASR Leaderboard at launch.
