Audio Transcriber
Turns an audio recording into timestamped text.
This tool runs on JavaScript. To use it, turn on JavaScript in your browser and reload the page.
TURN THE RECORDING INTO TEXT.
Audio Transcriber turns the speech in an audio or video recording into text. There are three outputs: plain text gathers the speech into paragraphs at sentence endings and suits turning it straight into writing; timestamped text puts the time at the start of each line in the form [00:01:23], which is handy for meeting notes and interview transcripts; JSON gives the start and end time of every word, for use in your own application or in editing. For speech recognition, OpenAI's open Whisper model runs on the server; 16 languages including Turkish can be selected, or it is detected automatically.
You upload the recording only from your own device; there is no link field. The audio is separated from the file, reduced to 16 kHz mono and passed to the Whisper model on the server; it never goes to an outside service and is deleted when the job ends. Parts without speech (silences longer than half a second, waiting) are filtered out before transcription. There are three models: “Fast” gives a rough draft; “Balanced” is enough for most recordings; “Best” catches names more accurately but is 3–4 times slower and works on recordings up to 10 minutes. A recording can be at most 30 minutes; split a longer lecture or meeting into parts with “Audio Cutter” and upload them one by one. Pick the language if you know it; if automatic detection is unsure, the result screen tells you.
From start to finish
- Upload your audio or video file. MP3, WAV, M4A, FLAC, MP4, MOV, MKV… The duration is read.
- Choose the language and output. Plain text, timestamped text or JSON. Model: Balanced; Best if names matter.
- Download the text. The text appears on screen and can be copied; the file downloads. Read and correct.
Frequently asked questions
What is timestamped text for?
The time at the start of each line shows where in the recording that sentence occurs; it is used for meeting notes, interview transcripts and finding a passage inside a video. The JSON output gives the same information word by word.
Does it separate speakers?
No. The text comes as a single stream; who is speaking is not marked. In a two-person interview, the quickest way is to choose “Timestamped text” and add names at the start of lines while listening to the recording.
How are paragraphs formed in plain text?
The model transcribes speech in segments; the plain-text output joins them and starts a new paragraph at the first sentence end after about 400 characters. Punctuation and capital letters also come from the model and can be missing in fast or noisy speech. Read the text once before publishing and check proper names in particular.
Does my file stay on the server?
No. Your file reaches the server only to be processed and is deleted when the job is done; the output is deleted too once you download it and leave the tool. You don't need an account, a name or an email address, and your files are never sent to a third-party service. It is free to use: we set no hourly or daily quota on the number of jobs, we don't make you watch an ad before processing, and we add no watermark to the output.
Other Audio tools
RUPO Studio