100% local — your file never leaves your browser
Subtitle Generator
Generate timed subtitles from any video or audio file — free, in your browser, with nothing uploaded and no sign-up.
- Free
- No sign-up
- No watermark
How it works
- 01
Add your video or audio
Pick the file or drag it onto the page. The audio track is read directly from your disk — the picture is never decoded and nothing is sent anywhere.
- 02
Pick the spoken language
Choose the language rather than letting it be guessed. Timing comes out of the same pass, so cues are aligned to speech automatically.
- 03
Export SRT or VTT
Download SRT for a video editor or a social platform, VTT for HTML5 video on your own site. Plain TXT is there too if you only want the words.
Why choose TranscriptSnap
Cues, not just a wall of text
Every line comes with a start and end timecode aligned to the speech, so what you download is a working subtitle file rather than a transcript you would still have to time by hand.
SRT and VTT, both from the same run
SRT for video editors and social platforms, VTT for HTML5 video on your own site — that second case will silently ignore an SRT, which is the one hard rule in this format choice.
Subtitling happens before publication
Which is exactly when a video is most sensitive: a rough cut, an embargoed interview, a client edit under review. The file is read from your disk in this tab and never sent anywhere.
The text stays editable
You get a sidecar file, not burned-in pixels. That means you can shorten the long cues, fix a name, or translate the whole thing later without going back to the video.
Good to know
What you are actually generating
A subtitle file is a list of cues, and a cue is two things: a start and end timecode, and the line or two of text that should be on screen between them. That is the entire format. Everything that makes subtitles good or bad happens in how those cues are cut — where one ends, how much text is inside it, and whether a viewer can finish reading before it disappears. The transcription is the easy half; the cue timing is the half that decides whether anyone can actually read them.
SRT or VTT
Both do the same job and differ in small details — a comma versus a period before the milliseconds, a header line, optional cue numbering. The one hard rule is that HTML5 video accepts only VTT, so subtitles for your own website must be VTT or the browser silently ignores them. For a video editor, a social upload or a desktop player, SRT is accepted almost everywhere and is the safer default when you are unsure. The practical difference between SRT and VTT
Why generated cues usually need shortening
Automatic cues follow speech, and speech is faster than reading. A sentence that took four seconds to say frequently cannot be read in four seconds, especially by someone half-watching on a phone. Broadcast practice puts comfortable reading at roughly fifteen to seventeen characters per second, with a maximum of two lines and about forty characters per line. When a generated cue busts that, the fix is not to extend its timing — that desynchronises everything after it — but to cut words. Subtitles are allowed to be a tightened version of what was said, and the good ones usually are.
Sidecar file or burned in
A sidecar is the subtitle file sitting next to the video, switched on by the player: editable later, translatable, and skippable by viewers who do not want it. Burned-in means the text is rendered into the picture during export, so it always shows — which is why social platforms are dominated by burned-in captions, since much of the audience watches muted and never touches a settings menu. Generate the file either way; where it ends up is an export decision you make afterwards. How to put a transcript back onto a video as subtitles
The file never leaves your browser
Subtitling usually happens before publication, which is exactly when a video is most sensitive — a rough cut, an embargoed interview, a client edit under review. The speech model here runs inside the browser tab, so the file is read from your disk and stays on it. No account, no upload, no retention window, and you can verify that in the network panel while it runs rather than taking it on trust.
Where the cues will be wrong
Timing is generally reliable; wording is where to spend your editing time. Proper nouns and product names are the largest cluster of errors, because the model has nothing to check them against. Overlapping speakers merge into a single run of text with no indication that two people were talking. Music mixed on top of dialogue costs more accuracy than any other single factor. And the model never leaves a blank when unsure — it writes the most plausible word instead, so a wrong subtitle looks exactly as confident as a right one. Why transcription tools invent words that were never said
Frequently asked questions
Is this subtitle generator free?
Yes, and unlimited. Everything runs on your own device, so there is no per-minute cost to meter and no daily cap.
Which subtitle formats can I export?
SRT and VTT, both with timestamps, plus plain TXT if you only need the words.
Should I use SRT or VTT?
SRT for video editors and social platforms. VTT if the video will play in an HTML5 player on your own site — that case requires VTT.
Is my video uploaded?
No. It is read from disk by your browser and processed in the same tab, which you can confirm in the browser network panel.
Can it burn the subtitles into the video?
No. It produces the subtitle file; burning in happens in a video editor at export time. That also keeps the text editable for as long as possible.
Will the cues be the right length?
They follow speech, which is usually a little too fast to read comfortably. Expect to shorten the longest cues by cutting words rather than extending their timing.
Does it label who is speaking?
No. Speaker labelling needs a second model that is not practical in a browser, so overlapping speakers appear as one continuous stream.
Can I subtitle a file that is only audio?
Yes. Audio files work exactly the same way and still produce timed cues — useful for adding captions to an audiogram later.
A free subtitle generator that produces SRT and VTT locally
Generating subtitles is two jobs that get treated as one. The first is recognising the speech, which modern models do well. The second is cutting it into cues a person can actually read in the time each one is on screen, which is where automatically generated subtitles usually fall short — speech is faster than reading, and machine cues follow speech. This subtitle generator handles the first job entirely in your browser, with nothing uploaded and no limit on length, and gives you a standard SRT or VTT file for the second. Expect to shorten the longest cues by cutting words rather than stretching their timing, and you will end up with subtitles that read the way good ones do.