Whisper is an open-source speech recognition model released by OpenAI in 2022. One model transcribes many languages, detects which language is being spoken and can translate speech into English. It was trained on 680,000 hours of audio from the internet. Accuracy varies by language, accent and recording quality, and it can produce text nobody said, especially over silence.
Most apps that turn speech into text today are built on a small number of speech models, and Whisper is the one you are most likely to meet. Knowing a little about it explains a lot of everyday behavior: why one recording comes back nearly perfect and another does not, why a name is wrong every time, and why a silent stretch sometimes gains a sentence.
Everything below comes from three primary sources: the paper Robust Speech Recognition via Large-Scale Weak Supervision, the GitHub repository with its model card, and the large-v3 model card on Hugging Face.
What is Whisper?
Whisper is a speech recognition model from OpenAI. The repository describes it as “a general-purpose speech recognition model” that can perform multilingual speech recognition, speech translation and language identification. The code and model weights are released under the MIT License.
“Open source” matters here in a practical sense. Anyone can download the model and run it on their own computer or server. No account with OpenAI is needed, and no audio has to be sent to OpenAI to use it.
How was Whisper trained?
Earlier speech systems were usually trained on a modest amount of carefully transcribed audio. Whisper took the opposite route: a very large amount of imperfect data. The paper’s abstract describes training on “680,000 hours of multilingual and multitask supervision”, meaning audio from the internet paired with transcripts that already existed for it.
The model card breaks that down:
| Share of the training data | Hours | What it is |
|---|---|---|
| About 65% | 438,000 | English audio with English transcripts |
| About 18% | 126,000 | Non-English audio with English transcripts (translation) |
| About 17% | 117,000 | Non-English audio with transcripts in the same language |
The non-English portion covers 98 languages, according to the model card.
The authors call this “weak supervision” because the transcripts were not checked by hand. The benefit is robustness: the model heard many accents, microphones and background noises. The cost is that it also learned from mistakes in the data, which comes up again below.
What is large-v3?
Whisper comes in several sizes, from tiny to large. Bigger models are slower and more accurate. Large-v3 is the third version of the largest one. The model card gives its release date as November 2023.
Per the large-v3 model card:
- It has the same architecture as the earlier large models, with two changes: the input uses 128 Mel frequency bins instead of 80, and there is a new language token for Cantonese.
- It was trained on 1 million hours of weakly labeled audio and 4 million hours of pseudo-labeled audio, the latter labeled using Whisper large-v2.
- It shows a 10% to 20% reduction of errors compared with large-v2 over a wide variety of languages.
That last figure is relative. It means fewer errors than the previous version, not a particular accuracy on your recording.
How does language detection work?
Whisper identifies the language as part of the same model that transcribes. The paper describes it as the first thing the model predicts, using one token per language in its training set, 99 in total. In the reference implementation, when no language is given, the code reports that it is “detecting language using up to the first 30 seconds” (transcribe.py). The repository’s README also shows that you can skip detection by naming the language yourself.
Detection is good but not perfect. The paper’s own evaluation says Whisper’s language identification “is not competitive with prior supervised work” on the Fleurs benchmark, partly because the training data had no examples at all for 20 of that benchmark’s 102 languages. A wrong guess at the start affects everything after it, which is why many apps let you set the language by hand.
Why does accuracy vary by language?
The main reason is training data. The paper found a strong relationship between how many hours of a language the model saw and how well it transcribes that language, and estimates that word error rate “halves for every 16× increase in training data”. Languages with little audio on the internet get worse results.
The model card says the same thing in plain terms: the models “perform unevenly across languages”, with “lower accuracy on low-resource and/or low-discoverability languages”. It adds that the models show strong results in about 10 languages. The README is blunter: “Whisper’s performance varies widely depending on the language.”
Why do accent, noise and recording quality matter?
- Accents and dialects. The model card says the models “exhibit disparate performance on different accents and dialects of particular languages”, which may include higher error rates for speakers of different genders, races or ages.
- Noise. The paper tested Whisper with added noise and found it more robust than the models it was compared with, especially under “pub noise”, a recording of a noisy room. Robust is not immune. Every model in that test got worse as noise increased.
- Recording quality. A quiet voice far from the microphone leaves the model less to work with. Distance, echo and people talking over each other all make transcription harder, whatever the model.
The first thing to fix is usually the recording itself: put the device closer to the people speaking.
What about switching languages mid-sentence?
Many people do this every day. Whisper handles it poorly. In a discussion on the repository, a maintainer wrote that the model is “intended for monolingual audio inputs” and that it “doesn’t support code-switching inputs very well”. Part of the reason is in the training process: the paper explains that audio whose spoken language did not match the transcript’s language was filtered out of the transcription data.
In practice, a mixed-language recording may come back mostly in one language, with the other parts translated, garbled or missing.
What are the known failure modes?
The authors are open about these.
- Hallucination. The model card says predictions “may include texts that are not actually spoken in the audio input”. The suggested cause is that the model combines predicting the next word with transcribing the audio.
- Repetition. The same document says the architecture “makes it prone to generating repetitive texts”.
- Long recordings. Whisper works on 30-second pieces of audio. For long-form transcription the paper lists “getting stuck in repeat loops, not transcribing the first or last few words of an audio segment, or complete hallucination where the model will output a transcript entirely unrelated to the actual audio”.
- Silence. The model is trained to output a special no-speech token when a segment contains no speech, but the paper found that this signal “alone is not sufficient to distinguish a segment with no speech”. Extra thresholds make the check “more reliable”. When it fails, a silent stretch gets text that nobody said.
The model card adds that hallucination and repetition are likely worse in lower-resource languages.
What is word error rate?
Word error rate (WER) is how speech recognition is scored. Take a correct reference transcript and compare the model’s output to it:
WER = (substituted words + deleted words + inserted words) ÷ words in the reference
If the reference has 100 words and the output gets 5 wrong, drops 3 and adds 2, the WER is 10%. Lower is better. For some languages the README reports character error rate (CER) instead, which counts characters rather than words.
A WER figure describes one model on one dataset. A number measured on audiobooks read in a studio tells you little about a meeting recorded across a table. That is why a single accuracy percentage for a product is rarely meaningful.
Why does running an open model on your own servers matter?
A company that wants transcription has two options. It can send your audio to another company’s speech service, or it can run a model like Whisper on machines it controls. With the first, a third party receives the recording and its terms apply to it. With the second, the audio goes to one place, and the operator decides what is kept and for how long.
Self-hosting is not a guarantee of anything by itself. What matters is what the operator does: whether audio is deleted, whether it is used for training, who can access it. Those are the questions to ask of any tool, and AI note taker privacy lists them.
The model card also carries a caution that applies to everyone who builds on Whisper: it warns against using the models “to transcribe recordings of individuals taken without their consent”. Recording other people may require their consent depending on where you are. Tell people before you record.
How does Nuvi use Whisper?
Nuvi, the AI note taker for iPhone, iPad and Apple Watch, uses Whisper large-v3, run on servers the developer operates. Audio is not sent to a third-party AI service. It is uploaded over an encrypted connection, transcribed, and the upload is deleted once the transcript has been returned. The privacy-first page has the full description.
Nuvi detects the spoken language by itself, and you can pin one language under Settings > Spoken language if detection gets it wrong. Transcription and summaries work in many languages, including English, Spanish, French, German, Italian, Portuguese, Dutch, Turkish, Polish, Russian, Ukrainian, Arabic, Hindi, Japanese, Korean and Chinese. Quality is not equal across them, for the reasons above, and Nuvi publishes no accuracy percentage.
The limits of the model are the limits of the app. Transcripts can be wrong or incomplete. A brand name said the local way may be misheard; “Fix a word…” corrects it in that note and in later recordings, and a custom vocabulary helps with names and jargon. Telling speakers apart is a different problem, and Whisper’s model card says the models have not been robustly evaluated for it; speaker diarization explained covers that topic.
The app’s interface is in English, and transcription needs an internet connection. Get Nuvi. The full list of what it does is on the features page.