Subtitle Generator
Transcribes speech and creates a subtitle file.
This tool runs on JavaScript. To use it, turn on JavaScript in your browser and reload the page.
TURN SPEECH INTO SUBTITLES.
Subtitle Generator turns the speech in a video or audio file into text and gives it to you as a timed subtitle file. SRT output is used in players, in YouTube uploads and in editing programs; VTT output is used in web players (HTML5). For speech recognition, OpenAI's open Whisper model runs on the server; 16 languages including Turkish can be selected, or the language is detected automatically. Subtitle cues are rebuilt to stay readable: at most two lines, at most 42 characters per line, at most 6 seconds, cut at sentence endings.
The recording comes only from the file you upload; there is no field for pasting a link or a video address. The audio is transcribed with a Whisper model running on the server; it is never sent to an outside service and is deleted when the job ends. Pick the language if you know it: automatic detection looks at the first 30 seconds and can be wrong on mixed or noisy recordings. The “Balanced” model is enough for most recordings and the job takes about as long as the recording; the “Best” model catches names and jargon more accurately but is 3–4 times slower and is used for recordings up to 10 minutes. A recording can be at most 30 minutes; if it is longer, split it first with “Video Cutter” or “Audio Cutter”. The result appears on screen, can be copied and downloads as a file; always read and correct it before publishing — proper names, jargon and noisy passages are where the most mistakes happen.
From start to finish
- Upload your video or audio file. MP4, MOV, MKV, WebM, MP3, WAV, M4A, FLAC… The duration is read.
- Choose the language and format. Pick the language if you know it; SRT for most uses, VTT for web players. Model: Balanced.
- Download the subtitles. The cues appear on screen and can be copied; the file downloads. Read and correct, then publish.
Frequently asked questions
SRT or VTT?
SRT is the most common format: VLC, YouTube, Premiere, DaVinci Resolve and most players read it. VTT is the format the HTML5 video tag in the browser expects; choose VTT if you will embed it in a website. The content is the same, only the notation differs.
Can I give a YouTube or other link?
No. The tool works only with the file you upload; there is deliberately no field for pasting a link. Downloading another service's content is against that service's terms of use.
How accurate is it?
On a clean recording (one speaker, close microphone) the “Balanced” model gets the great majority of words right; proper names, abbreviations and technical terms can get mixed up. Noise, a distant microphone and people talking over each other raise the error rate. The “Best” model is noticeably more accurate but slower. Treat the output as a draft and read it before publishing.
Does my file stay on the server?
No. Your file reaches the server only to be processed and is deleted when the job is done; the output is deleted too once you download it and leave the tool. You don't need an account, a name or an email address, and your files are never sent to a third-party service. It is free to use: we set no hourly or daily quota on the number of jobs, we don't make you watch an ad before processing, and we add no watermark to the output.
Other Audio tools
RUPO Studio