Captions
Captions for client videos:
generate the draft, fix the names, deliver the file.
A speech-to-text model can write a usable transcript of an interview in the time it takes to make coffee. It cannot spell your client's product, it cannot punctuate a room with three people talking, and it does not know which of them is worth naming on screen. The machine does the typing. You do the pass that makes it deliverable. Here is the whole territory: what breaks, what the client should actually receive, and when a generated track is the wrong answer entirely.
Video hosting for filmmakers · 5 GB free · Paid plans from USD 9/month
Updated September 2026
Why it is on the list
Most of the audience arrives while
the sound is off.
Captions stopped being an accessibility line item and became a delivery format, because the places client work ends up mostly start playback silent. The text is not a courtesy on top of the film. On a feed it is the film, for the first three seconds, and if it is not there the viewer scrolls past a moving picture they never heard.
Feeds that autoplay muted
Social and in-page players start silent by default and wait for a tap that most viewers never give. A cut with no on-screen text has to carry its whole meaning in pictures. A cut with text keeps working in a train carriage, an open-plan office and a waiting room.
Accessibility obligations
Public sector, healthcare, education and broadcast clients frequently carry a captioning requirement in the contract, not in the brief. Ask which standard applies before you deliver, because the answer decides whether a generated track is acceptable at all or whether the job needs a certified service.
Viewers reading a language they do not hear well
A second-language viewer often reads a language far more comfortably than they follow it spoken over room tone and music. Same-language captions widen the audience for a corporate film more than a translation budget usually does.
There is a fourth reason that has nothing to do with the audience. A caption track is a timecoded transcript, which makes the dialogue of a film searchable. That turns a two-hour interview from something you scrub through into something you query, and it is worth its own page: searching the dialogue inside a video covers how the caption panel here is used as a finding tool rather than a reading one.
Get the word right first
Captions and subtitles are not the same thing,
and the client means one of them.
Captions are written for a viewer who cannot hear the audio. They carry speaker labels when it is not obvious who is talking, and they carry non-speech sound that changes the meaning of a scene: a door, a laugh, music starting. Subtitles are written for a viewer who hears everything and does not follow the language, so they translate the dialogue and leave the rest of the sound design alone.
The distinction matters commercially, not academically. When a client asks for subtitles they usually mean same-language captions burned or attached, which is a proofreading job. Occasionally they mean an actual translation into another language, which is a translation job, priced by the word, done by a person who speaks it, and reviewed by somebody at the client who also speaks it. Those two quotes are not close to each other. One question at the briefing stage prevents the awkward version of that conversation after delivery.
One more term you will meet in a brief: SDH, subtitles for the deaf and hard of hearing. It is the hybrid. Translated or same-language text that also carries the speaker labels and the sound cues. If a distributor asks for SDH, they are asking for captioning discipline in a subtitle file, and they usually have a style guide to go with it.
The part nobody can skip
Transcription is accurate right up until it lands
on the words your client is paying for.
A speech-to-text model predicts the most likely next word from everything it has heard before. That is why it is excellent on ordinary sentences and consistently wrong in the same five places, all of which happen to be the places a client reads first.
Proper nouns
Names of people are the single biggest failure. A model will render a surname as the nearest common word, and it will do it consistently through a whole interview, so one wrong guess becomes forty wrong lines. Every name that appears on a lower third has to be checked against the lower third.
Brand and product names
Invented names are, by definition, not in ordinary language, so they come back as the closest real words. A product spelled with a dropped vowel or an internal capital is guaranteed to be wrong. This is the error that gets noticed, because it is the client's own name spelled wrong on their own film.
Jargon and acronyms
Industry vocabulary and three-letter acronyms are transcribed phonetically or expanded incorrectly. A trade term that the whole audience knows will come back as something that reads like nonsense to them and looks fine to you, because you do not work in that trade.
Accents and code-switching
Accuracy is uneven across accents, and it drops further when a speaker switches language mid-sentence, which is normal in a bilingual market. Regional vocabulary and place names go the same way. Read the speakers you found hardest to follow on set first, because the model found them hardest too.
Overlap, room and music
Two people talking over each other produces one merged line attributed to nobody. A loud room or music under dialogue produces dropped words rather than obvious gaps, which is worse, because a sentence that is missing two words still reads as a sentence.
Punctuation and line breaks
Punctuation is a guess about intent. A generated track will run two thoughts into one sentence, break a line mid-phrase, and hold a caption on screen through a cut. None of that is factually wrong and all of it reads badly, so a pass for rhythm is part of the job, not a polish step.
The reading order that finds the most errors fastest
The workflow
The machine drafts, the editor fixes,
the client gets a file.
Generate the first pass
A speech-to-text model transcribes the film and times the lines against it, so you never type a transcript from scratch. Generation is metered in minutes a month by plan, and the free Starter plan includes 20 minutes, which is enough to test the transcript quality on your own audio before you commit to anything.
Fix the names, then the rhythm
The owner edits the track directly. Correct every proper noun, brand name and piece of jargon against what is actually on screen, then fix the breaks that land mid-phrase. This is the pass that turns a transcript into a caption track, and on a well-recorded interview it is minutes rather than hours.
Let the review round catch what you cannot
A reviewer can submit a correction, and the creator applies it. Use it deliberately: the client knows their own product names, their staff surnames and the acronym their industry uses, and you do not. Corrections arrive attached to the track rather than as a list in an email you have to translate back into timecodes.
Download the file and deliver it
The finished track downloads as .srt or .vtt, and those two formats only. That is the file the client uploads with the video wherever it is going. A vaulted master refuses caption download, so if the text has to leave, caption the version you are allowed to circulate.
Two things follow from that fourth step. The client receives a video and a text file, not a video with text in it, which means the caption survives independently of the render. And the transcript exists as text, which is what makes the caption panel searchable: it has a search field, and every matching word is a button that seeks the player to that moment. For the mechanics of attaching the track, step by step, see adding subtitles to a client video.
The delivery decision
A file beside the video, or text
burned into the picture.
This is the question clients ask in the wrong direction. They ask which looks better. The question that decides it is what happens when somebody wants a word changed after the film has been published in four places.
Both can be true on the same job: deliver the sidecar file as the source of truth, and burn text into the social cutdowns that need it. The order matters. Approve the text as a caption file first, then burn from the approved text, so the two never disagree.
The default for client work is the sidecar file, because client work gets revised. The single change that costs an hour with burned-in text costs a minute with a caption file, and the file is the thing the client can hand to their web team, their agency and whoever cuts next year’s version. What you give up is typography, and that is a real loss on a piece where the text is design rather than access.
The exception worth naming is the vertical cutdown. Platform caption rendering is inconsistent, some publishing routes drop the track, and the type in those cuts is often doing creative work that a player’s default styling cannot do. Burn those. Then still deliver the caption file alongside, because the client will need the text again.
What lands in the client folder
Three things: the film, the file,
and the text that made it.
A caption deliverable is not one item. It is the film in the format the client asked for, the caption track as a .srt or .vtt beside it, and, where the client publishes text, the plain transcript that the track is built from. The third one costs you nothing because it already exists, and it is often the piece the client’s marketing team values most: it becomes the page copy, the pull quotes and the search-engine text under the embed.
Name the files so the pairing is obvious. A caption file whose name matches the video file except for the extension is picked up automatically by most players and by every editor who opens the folder six months from now. If you are delivering more than one language, the language code belongs in the filename, not in a note in the email. The rest of the handover, including what formats to ship and what to keep, is covered in delivering final video files to clients and in the deliverables checklist.
Where this is the wrong tool
A generated track is a first draft,
not a deliverable.
Say it plainly, because the whole category is marketed as though the machine finishes the job. It does not. If you are captioning legally sensitive material, broadcast content, or anything with a compliance standard written into the contract, you need a professional captioning service and a human quality-control pass, not a generated track with an editor’s read-through on top. That includes evidence, regulated financial and medical communication, anything a distributor will technically check, and anything where a mistranscribed sentence has a consequence beyond embarrassment.
The reason is not that generated text is bad. It is that those jobs are judged against a written standard covering reading rate, line length, placement, speaker identification and sound description, and they are signed off by someone who does this for a living. Read the contract before you decide. If a standard is named in it, quote the captioning service into the budget and stop treating it as a task you absorb.
For everything else, which is most corporate, commercial and social work, the honest description is the one at the top of this page. The model does the typing. You fix the names and the jargon. The client gets a caption track and a file they can hand anywhere else. That is a good workflow, and it is a much better one than either typing a transcript by hand or shipping a machine draft unread.
Questions
Frequently asked
What is the difference between captions and subtitles?
Captions are written for someone who cannot hear the audio, so they carry speaker labels and meaningful non-speech sound. Subtitles are written for someone who can hear the audio but does not follow the language, so they translate dialogue and leave the rest. In practice most clients say subtitles and mean captions in the same language as the audio. Ask which one they mean before you quote the job, because a translation pass is a different piece of work with a different price.
Are machine-generated captions good enough to deliver?
Not as they come out. Transcription is reliable on clean dialogue in a common accent and unreliable on exactly the words a client cares most about: their company name, their product names, their staff names and their industry jargon. Treat the generated track as a first draft that saves you the typing, then read it against the picture once and fix the names.
Should I burn captions into the picture or deliver a separate file?
Deliver a separate file for almost everything. A sidecar caption file can be edited in a text editor, turned off by the viewer, read by assistive technology and corrected without re-exporting the video. Burn text in only when the type is part of the design, or when you are delivering to a destination whose own caption rendering you do not trust. Some teams deliver both, which is why the source track matters.
What caption formats can I download here?
Two: .srt and .vtt. Those are the two the rest of the world accepts, so they cover uploading to a video platform, attaching to a web player, importing into an edit and handing a client something they can open. There is no burned-in export and no broadcast interchange format.
Can my client fix a caption without emailing me a list?
A reviewer can submit a correction against the track, and the creator applies it. That matters most for the words a reviewer knows and the model does not: internal product names, a colleague's surname, an acronym used only inside that company. The edit stays with the owner, so nothing changes on the film without you applying it.
How much caption generation do I get?
It is metered in minutes a month, and each plan carries its own allowance. The free Starter plan includes 20 minutes a month, which is enough to caption a short piece and see whether the transcript quality holds up on your kind of audio before you pay for anything.
Why will a vaulted master not give me a caption file?
A vaulted master refuses caption download by design. The vault exists so a sensitive cut cannot leave in a form you did not intend, and a caption file is a full transcript of everything said in it. If you need the text out, caption the version you are allowed to circulate rather than the vaulted one.
Caption a cut and hand over the file
Caption generation is metered in minutes a month by plan, and the free Starter plan includes 20 minutes. Edit the track, then download it as .srt or .vtt for wherever the client is publishing.
