About TranscriptSnap
TranscriptSnap turns spoken audio into text. Add an audio or video file from your device and you get back a transcript with timestamps that you can copy, or export as TXT, SRT or VTT. There is no account, no upload and no limit on how much you transcribe.
Who runs it
TranscriptSnap is built and operated by Novrik Digital LLC, a small software company registered in the State of Wyoming, United States. We make the product, write the guides on this site, and answer the email that comes in — there is no outsourced support desk between you and the people who wrote the code. You can reach us through the contact page.
How it works
The speech model — OpenAI's Whisper, converted to run with WebAssembly — runs inside your own browser. Your file is never uploaded, because there is no server to upload it to. The first run is slower while the model downloads and caches; after that it works offline, and there is no limit on how much you transcribe.
That design is the whole point rather than a technical curiosity. Interviews, medical notes, client recordings and unreleased material can be transcribed without any third party ever holding the audio — which does not just reduce the disclosure risk, it removes the disclosure entirely. You can check the claim yourself: load the page once, disconnect from the network, and transcribe again. How to verify a transcription tool never uploads your audio.
Whisper comes in several sizes. We ship two: a 73 MB model for languages it already handles well, and a 237 MB model for languages where the smaller one makes too many mistakes. Which one loads is decided by the language you pick, so an English speaker is not made to download three times more than they need.
Which languages, and why so few
The language menu lists sixteen languages, including English, Spanish, Portuguese, French, German, Russian, Vietnamese, Japanese, Korean and Chinese (written out in Simplified or Traditional characters). Whisper itself can attempt close to a hundred. We list far fewer on purpose.
Each language was measured before it was added. Some results were not what you would guess: Polish, a Latin-script European language, came out with roughly 30% of words wrong on the smaller model — worse than Russian — while Dutch was exactly as accurate on the smaller model as on the larger one. Hindi, Arabic, Turkish and Ukrainian were removed after testing, because even on the larger model between one word in six and one word in two was wrong. A transcript with that many errors still reads as fluent, confident prose, which makes the errors hard to notice. We would rather not offer a language than offer one that quietly produces a wrong record.
You choose the language rather than having it guessed. When automatic detection guesses wrong, it does not fail with an error message; it produces a plausible-looking transcript in the wrong language or invents text. Keeping the choice with you removes that failure entirely.
How accurate it is
On clear speech in a well-supported language, the transcript is usually close to word-for-word. Accuracy drops with background noise, music, several people talking over each other, heavy accents, and phone-quality recordings — the same things that make speech hard for a person to follow. The measurements above were taken on clean, synthetic speech, so they are a best case, not a guarantee.
Speech models have one failure worth knowing about: during long silences or unclear passages they can produce sentences that were never spoken. Why transcription tools invent words explains when that happens and how to catch it. For anything you will quote, publish or rely on legally, read the transcript against the recording.
What it deliberately does not do
It does not accept links. Pasting an Instagram or TikTok URL is not supported, and that is a decision rather than a missing feature. Resolving a post to its media file server-side requires third-party residential proxy networks, which break whenever a platform changes and add a dependency we would have to keep repairing. Taking a file instead is what lets "your audio never leaves your device" stay true for the whole product, with no asterisk.
It does not label speakers. Diarization needs a second model that is not practical to run in a browser, so the output is one continuous stream of text rather than a dialogue split by speaker. Adding speaker labels to a transcript covers the workarounds.
It does not keep anything. There is no history, no saved projects and no cloud library. When you close the tab, the transcript is gone unless you exported it.
Phones and large files
The tool works on phones, but a phone browser has much less memory for decoding video than a computer. Long videos, and MOV files from iPhones in particular, can fail to open on a phone. Exporting the audio as MP3 or M4A, trimming the file, or opening the page on a computer almost always solves it.
Why it is free, and how it is funded
The work happens on your device, so there is no per-minute server cost for us to charge back to you. That is why there is no account, no email, no watermark and no paid tier, and nothing is held back behind one. Our remaining costs — hosting the site and delivering the model files — are small, and we plan to cover them with advertising. Ads never receive your files or transcripts; the Privacy Policy explains exactly what advertising partners can and cannot see.
Guides
Alongside the tool we publish practical guides on transcription: how to transcribe a podcast, an interview or a lecture, the difference between SRT and VTT subtitles, how accurate Whisper is in different languages, and how to compare transcription tools on privacy. They are written from what we measured while building the product, not rewritten from other sites.
Contact
Questions, bug reports and suggestions are welcome at hello@transcriptsnap.com. The contact page lists what to include so a problem can be reproduced quickly.