MediaChef

MediaChef Video to SRT

Turn a video into SRT subtitles

Drop the video into MediaChef, pick «Make SRT subtitles for a video», leave the model on small and the language on auto, and run it. A .srt file lands next to the video with timings already in it. Everything happens on your machine: the speech never leaves the disk, and once the model is downloaded the recipe works with the network off. On an M5 laptop, 2 minutes 43 seconds of speech took 6.2 seconds with the default model — about 26 times faster than real time — and produced 73 subtitle cues averaging 39 characters, which is short enough to read comfortably.

What you need
MediaChef 0.8.5 plus a one-time model download
Default model
small — 488 MB, downloaded once and then kept
Speed
≈26× real time on the default model (measured, M5)
Works offline
Yes, after the model is on disk
Formats
SRT, VTT, plain TXT and JSON — one recipe each
What you get
clip.subs.srt next to the video, original untouched

Free and open source — read the code on GitHub All files and release notes

How to make subtitles for a video

  1. Download MediaChef

    One file for macOS, Windows or Linux. FFmpeg and the Whisper runner both travel inside the download — there is nothing to install separately and nothing to put on your PATH.

  2. Download a model, once

    The first transcription asks for a speech model. small is the default at 488 MB and the one these measurements use; tiny is 78 MB, base 148 MB, large-v3-turbo 1.62 GB. It is fetched once, stays on disk, and after that the recipe never touches the network again.

  3. Drop the video in and pick the recipe

    «Make SRT subtitles for a video» takes the video straight — you do not need to pull the audio out first. MediaChef decodes the sound to the 16 kHz mono that Whisper wants, in a temporary folder you never see.

  4. Run it and open the .srt

    The file lands next to the video as clip.subs.srt, with numbered cues and timestamps. Players, editors and video platforms all read it directly, and it is plain text, so you can fix a name or a term in any editor.

MediaChef ready to convert: the board waits for a video file, the job queue is on the right.
The board the video lands on. Recipes appear once MediaChef has read the file.

Which model to pick

Four models, the same 2 minutes 43 seconds of speech, the same machine — an M5 laptop with 16 GB, each model warmed up first and the best of two runs taken.

ModelDownloadTimevs real timeWords wrong
tiny78 MB2.1 s×785 of 540
base148 MB2.6 s×633 of 540
small — the default488 MB6.2 s×260 of 540
large-v3-turbo1.62 GB11.5 s×141 of 540

Read the last column carefully, because the test audio is a synthesised voice reading a prepared script: no accent, no background noise, nobody talking over anybody. That is why even the smallest model is almost perfect here, and it is not what a real meeting recording sounds like — on hard audio the gap between these models widens a lot. The speed column is the one that transfers directly. One more thing the raw comparison hid: most of the mismatches were numbers written as digits rather than words — large-v3-turbo wrote «70», «10», «50», «30» where the script said them in full — which is formatting, not mishearing.

How long the subtitle lines come out

A subtitle that is technically correct can still be unusable if it puts twenty words on screen at once. The models chop the same speech very differently, and this is measured on the same run as above.

ModelCuesAverage lengthAverage charactersLongest
tiny354.7 s8397 characters
base354.7 s83100 characters
small — the default732.2 s3958 characters
large-v3-turbo305.4 s97112 characters

The common broadcast guideline is about 42 characters a line over two lines, so roughly 84 characters on screen at once. By that measure small is the only one of the four that comfortably fits, at 39 characters on average and 58 at its longest, while large-v3-turbo runs past the limit on a typical cue. So the default model is not just the balanced choice on accuracy — it also chops the speech into the most readable pieces.

SRT, VTT, plain text or JSON

The same transcript written four ways. Sizes are from the same 2 minutes 43 seconds of speech, so they compare directly.

FormatSizeWhat is insideUse it when
SRT5.5 KBNumbered cues, timings with a comma: 00:00:00,000Almost always. Players, editors and platforms all take it
VTT5.3 KBA WEBVTT header, timings with a dot: 00:00:00.000Subtitles for a web player, a browser video track
TXT3.0 KBPlain running text, no timings at allYou want the words, not the subtitles
JSON15.2 KBEvery cue plus the model and parameters usedSomething else is going to read this, not a person

SRT and VTT differ mainly in the character between seconds and milliseconds, so if a player rejects one, the other is a one-recipe fix rather than a re-transcription. JSON is roughly three times the size of SRT because it carries the run's metadata alongside the text.

Which recipe for which job

Subtitles are not one recipe but several, and picking the right one saves a step. All of them are in the 17-recipe catalogue.

What you haveWhat you wantRecipe
A videoSubtitles next to itMake SRT subtitles for a video
An audio fileSubtitlesTranscribe audio to SRT subtitles
Speech in another languageEnglish subtitles, one passTranslate speech to English subtitles
Anything with speechJust the textTranscribe audio to text
Anything with speechA web player trackTranscribe audio to WebVTT

The translation recipe goes straight from foreign speech to timed English subtitles in a single pass — you do not transcribe first and translate after. It only goes to English, though; that is a limit of the model, not of the app.

Why make subtitles on your own machine

The speech never leaves your disk. Recordings of meetings, interviews and calls are the single most sensitive kind of file most people handle, and an online transcriber is by definition a copy of that conversation on somebody else's server. Here there is no upload to reason about.

No per-minute pricing. Transcription services bill by the minute of audio, which turns a long archive into a real invoice. The model download is a one-off, and after that a two-hour recording costs the same as a two-minute one: nothing.

It runs with the network off. Once the model file is on disk, nothing about this recipe touches the internet. It works on a plane, on a locked-down machine, and in a room where the wifi is the least reliable thing present.

No length limit. Free web transcribers usually cap you at a few minutes per file, exactly when a recording is worth transcribing because it is long. There is no cap here.

A whole folder at once. Drop in a directory of recordings and the queue works through them one at a time, telling you where each subtitle file was written.

When this is the wrong tool

The recipe writes a subtitle file. That is a narrower job than «putting subtitles on a video», and the difference matters in these cases.

  1. You want the subtitles burnt into the picture.

    This produces a separate .srt file that a player loads alongside the video. Burning the text permanently into the frames is a different operation — it re-encodes the video, and the words then cannot be turned off or edited.

  2. You need broadcast-grade accuracy.

    Even on the clean audio measured above the models slipped on a few words, and real recordings are harder. Anything published under a legal accessibility requirement gets read by a human before it ships, whatever produced the first draft.

  3. The audio is genuinely bad.

    Heavy crosstalk, a phone recording of a room, or music louder than the voice will defeat all four models. Fixing the audio first — even just extracting a cleaner track — does more for the result than moving up a model size.

  4. You need translation into something other than English.

    Whisper translates into English and only English. For any other target language, transcribe in the original language first and translate the resulting text with a tool built for that.

Questions

Is this free?

Yes, all of it. MediaChef is open source under GPL-3.0, there is no paid tier, no per-minute charge and no cap on file length. The models are free downloads too. Version 0.8.5 is the current one.

Does my video get uploaded anywhere?

No. The speech is processed by a model file on your own disk. The only thing that ever crosses the network is the one-time model download, and after that the recipe runs with the network off.

How long does it take?

About 26 times faster than real time on the default model: we measured 6.2 seconds for 2 minutes 43 seconds of speech on an M5 laptop. By that ratio an hour-long recording takes a couple of minutes. tiny ran at ×78 and large-v3-turbo at ×14 on the same audio.

Which model should I choose?

Start with small, the default. In our measurements it got every word of the test audio right and produced the most readable subtitle lines — 39 characters on average against 97 for large-v3-turbo. Move up only if your audio is difficult; move down to tiny or base if you want a rough draft in a couple of seconds.

How big is the model download?

78 MB for tiny, 148 MB for base, 488 MB for small, 1.62 GB for large-v3-turbo. It happens once. After that the file sits on your disk and every later run uses it without asking.

Do I need to tell it what language the speech is in?

No. Language is set to auto by default and the model works it out from the audio. You can still name the language explicitly, which is worth doing when a recording opens with a few words in another language.

Can it translate the subtitles into English?

Yes, with the «Translate speech to English subtitles» recipe: foreign speech in, timed English SRT out, in one pass rather than transcribe-then-translate. English is the only target language the model supports.

What is the difference between SRT and VTT?

Mostly the punctuation in the timestamps: SRT writes 00:00:00,000 with a comma and numbers its cues, VTT writes 00:00:00.000 with a dot and opens with a WEBVTT line. SRT is what players and editors expect; VTT is what a web player wants for its own subtitle track. Both come from separate recipes, so switching is a re-run, not a rewrite.

Can I edit the subtitles afterwards?

Yes — an .srt is plain text. Open it in any editor to fix a proper noun, a piece of jargon or a timing. This is the normal way to work: let the model do the ninety-something percent and correct the rest by hand.

Why are some of my subtitle lines too long?

Because the model decides where to break, and the bigger models break less often. We measured 39 characters per cue on small against 97 on large-v3-turbo, on the same audio. If your lines are running long, moving down to small usually fixes it, and it costs nothing in accuracy on clean speech.

Can it tell speakers apart?

No. Whisper writes what was said, not who said it. If you need «Speaker 1 / Speaker 2» labels, you will be adding them by hand or using a tool built specifically for that.

What happens if there is no speech in the file?

The run stops and tells you it heard nothing recognisable, rather than quietly writing an empty file. Silence produces no subtitles by design.

Does it work on Windows and Linux?

On all three platforms. Speech runs on the CPU everywhere and additionally uses the GPU on Apple Silicon, which is why the numbers above are fast — the same recipe on a modest Windows laptop will be slower, though still faster than listening to the recording.

Can I subtitle several files at once?

Yes. Drop in a whole folder, add the recipe, and the queue works through them one after another. Each subtitle file is written next to its own source.

Does the video file get changed?

No. A separate .srt is written next to it — clip.subs.srt — and the video is not modified, renamed or re-encoded. Nothing about this recipe touches the picture.

Get subtitles for that video

MediaChef 0.8.5 — free, open source, macOS · Windows · Linux.

Free and open source — read the code on GitHub All files and release notes

MediaChef is young: builds are not yet signed by Apple or Microsoft, so the first launch asks for confirmation — a plain-text how-to ships inside every download.

Something broken, or something missing? Write to us: hello@mediachef.app