OpenAI Whisper is a speech recognition model. You give it audio, it gives you text. It is the model most transcription services quietly run underneath, including several that charge by the minute for it.
The usual complaint is that using it yourself means installing Python, then ffmpeg, then PyTorch, then finding out your laptop has no CUDA. This guide skips all of that: Whisper can run directly in a browser tab.
What Whisper actually is
Whisper was trained on around 680,000 hours of multilingual audio. That scale is why it handles accents, background noise and code-switching better than the speech engines that came before it. It ships in several sizes, and the size you pick is the main quality/speed tradeoff you control.
| Model | Parameters | Rough size (quantized) | Where it makes sense |
|---|---|---|---|
| tiny | 39M | ~15 MB | Fast drafts, clean audio, weak devices |
| base | 74M | ~73 MB | English, Spanish, French, German, Dutch |
| small | 244M | ~237 MB | Everything else — and it is not optional |
| medium / large | 769M–1.5B | 1–3 GB | Server-side only in practice |
The usual advice is “base is the sweet spot for browsers.” That is true for English and it is badly wrong for most other languages. Measured word error rates on the same recorded content:
| Language | base | small |
|---|---|---|
| Spanish | 0.0% | 0.0% |
| French | 0.0% | 0.0% |
| Dutch | 3.7% | 3.7% |
| German | 4.2% | 0.0% |
| Portuguese | 6.7% | 4.4% |
| Korean | 8.5% | 4.3% |
| Chinese | 11.2% | 0.0% |
| Japanese | 11.5% | 1.6% |
| Vietnamese | 16.0% | 2.0% |
| Russian | 24.1% | 3.4% |
| Polish | 30.0% | 0.0% |
| Swedish | 45.0% | 5.0% |
| Turkish | 60.0% | 26.7% |
Polish is the one to remember. It is a Latin-script European language, so intuition
files it next to Spanish and French — and it scores 30%, the same league as Russian.
Swedish is worse still: 45% on base, 5% on small. You cannot predict this from
the language family or the writing system. Either measure it, or default to small
for anything that is not English.
The practical consequence: a browser tool that uses base for everything is quietly
serving broken transcripts to most of the world, and the transcripts read fine.
Running it without installing anything
Whisper has been compiled to WebAssembly, which means the model can execute inside the browser’s own sandbox. Nothing is uploaded — the audio is decoded locally, fed to the model locally, and the text comes back locally.
The practical consequences:
- The first run is slow. The model has to download once (about 74 MB for
base). After that it is cached and starts immediately. - Later runs work offline. Once cached, you can transcribe on a plane.
- There is no per-minute cost, because nobody is paying for a server to do the work.
- Long files depend on your RAM, not on someone’s upload limit.
If you want to try it right now, the audio to text page does exactly this — pick a file, and the model loads on first use.
You have to tell it the language — and here is why that matters
Whisper the model can detect language. Several Whisper implementations do not, and the difference is not academic.
The most widely used browser runtime for Whisper (transformers.js) has this in its
source, in the function that assembles the decoder’s opening tokens:
if (!language) {
// TODO: Implement language detection
language = 'en';
}
If no language is supplied it does not guess — it hard-codes English. Feed it Mandarin and the model is instructed to produce English from Chinese audio. It does not error. It does not return empty. It invents fluent English:
Spoken (Mandarin): 快速快放。今天的会议记录需要整理成文字。 Output: “The fast-food delivery of the product is now available for the customer.”
That output is not a translation — a translation would at least preserve the meaning. It is confabulation, and it is the single worst failure mode a transcription tool has, because nothing about the result looks wrong. No warning, no garbled characters, no confidence score. Just a clean English sentence that was never said.
So: if a transcription tool does not ask you what language the audio is in, find out what it does when you feed it something other than English before you trust it with anything that matters.
Two situations worth knowing about even when you do set the language:
- Heavy code-switching. One language token governs the whole file. A recording that alternates between two languages every few sentences will be forced into one of them.
- Chinese output defaults to Traditional characters. If you need Simplified, the conversion has to happen after transcription. Trying to steer it with an initial prompt costs you content — in testing, a prompt that successfully pushed the output to Simplified also dropped 32% of the words.
How accurate is it, really
On clean, single-speaker speech at a normal pace, base typically lands close to verbatim — good enough that editing is faster than typing from scratch. Accuracy drops in fairly predictable ways:
- Overlapping speakers. Whisper produces one text stream, so two people talking over each other blur together. It has no built-in speaker diarization.
- Proper nouns and jargon. Names, product names and technical terms are the most common errors. Whisper guesses a plausible-sounding word instead of leaving a gap, so these mistakes read as confident.
- Very noisy or distant audio. Phone recordings from across a room degrade sharply.
- Long silences. Older Whisper builds sometimes hallucinate text over silence; chunking the audio (see below) mostly avoids this.
That last point matters if you are transcribing something where accuracy is load-bearing. Whisper output is a strong first draft, not a court record.
Why chunking matters for long audio
Whisper processes 30-second windows. Naively cutting a long recording into 30-second slices splits sentences at the boundaries and loses words. The fix is overlapping windows — process 30 seconds, step forward 25, and reconcile the overlap.
Any tool worth using does this for you. If you ever see a transcript where words vanish at suspiciously regular intervals, that is what went wrong.
Whisper in the browser vs. the API
Both run the same family of models. The difference is where the work happens and what that costs you.
| Browser (WebAssembly) | Hosted API | |
|---|---|---|
| Your audio | Never leaves the device | Uploaded to a third party |
| Cost | Free, unlimited | Per minute of audio |
| First-run delay | ~74 MB download, once | None |
| Speed on long files | Limited by your CPU | Much faster |
| Works offline | Yes, after first load | No |
The honest rule: browser for anything sensitive or routine, API when you have hours of audio and need it done in minutes.
When you should not use Whisper at all
- You need speaker labels (“Speaker 1 / Speaker 2”). Whisper alone does not do this; you need a separate diarization step.
- You need certified or legally admissible transcripts. Those require a human transcriber, regardless of model quality.
- You need real-time captions during a live call. Whisper is designed for recorded audio; live captioning uses streaming models built for the job.
Common questions
Do I need a GPU? No. WebAssembly runs on the CPU. A GPU path exists via WebGPU but is not always faster or more reliable for quantized models — some combinations produce garbled output, so many tools deliberately stay on CPU.
Is MacWhisper the same thing? MacWhisper is a native macOS app wrapping the same model. If you live on a Mac and transcribe daily it is a fine purchase. If you want to transcribe one file without installing anything, a browser tool gets you there faster.
Can I transcribe a YouTube or Instagram video? Not directly from a link with a local-only tool — the model needs an audio file. Save the file you have the rights to use, then transcribe it.
Is the free browser version worse than the paid API? Not in model quality at the same size. The paid services are buying you speed and convenience, not smarter output.