Skip to content
uncompressed.io

Captions

Burned-in subtitles or an .srt file:
decide before the export, because only one of them is reversible.

Both put the same words on screen. One puts them inside the picture, permanently, at the moment you render. The other keeps them in a separate text file that anyone can open, correct, translate or switch off. The words are not the decision. Where the video is going is the decision, and getting it backwards costs a re-render, a re-upload and a second round of delivery emails.

See what each plan includes

Video hosting for filmmakers · 5 GB free · Paid plans from USD 9/month

Updated September 2026

Same words, two different objects

One of them is a picture,
the other is a file.

Everything on this page follows from that sentence. A burned-in caption is part of the image, with all the permanence an image has. A sidecar caption is a small text document that travels next to the video and gets drawn at playback by whatever is playing it.

Burned in

The text is rendered into the frame at export, so it is pixels. It always shows, it survives every player, every download, every re-post and every destination that quietly drops caption tracks. It also cannot be switched off, resized, repositioned, corrected, translated or read by anything that reads text. Changing one word means rendering the film again and re-delivering every copy already sent.

A sidecar file

The words live beside the video in their own file, with a timecode on every line. It can be opened in any text editor, fixed in seconds, swapped for another language, switched off by the viewer and read by assistive technology. What it depends on is a player that supports caption tracks, and a delivery route that keeps the two files together.

What neither one is

This is not a quality decision. Burned-in text is not sharper and a sidecar is not softer. What actually separates them is who controls the words after delivery: you, permanently, at render time, or the viewer, the client and whoever opens the file next year.

There is a formal name for what burning in does. The Web Content Accessibility Guidelines call it an image of text. Success criterion 1.4.5 Images of Text, at level AA, asks that text be used to convey information rather than images of text, with two exceptions: when the image of text can be visually customised to the user’s requirements, and when a particular presentation of the text is essential to the information being conveyed. Type that is part of a title card can reasonably be called essential. Dialogue rarely can.

Captions themselves sit lower down the same document, at the floor rather than the aspiration: success criterion 1.2.2 Captions (Prerecorded) is level A, and it asks for captions on all prerecorded audio content in synchronised media. Whether an open, burned-in caption satisfies an audit is a question for whoever signs the audit off. What is not in dispute is that a text track is the machine-readable option and a burn-in is not.

The other half of the folklore is that a sidecar means no typography at all. That is true of SubRip: an .srt file is a numbered list of timecodes and lines, which is precisely why everything reads it. WebVTT is less limited. Its specification defines cue settings for vertical text, line, position, size and alignment, permits STYLE blocks inside the file itself, and allows external CSS through the ::cue pseudo-element. What you cannot rely on is a given player honouring any of it. Plan for SubRip’s capabilities, and treat WebVTT styling as a bonus on players you control.

The only question that decides it

Where is this video going? Answer that and
the format answers itself.

Not what the captions should look like. Not which one is more professional. Each destination has one right answer, set by what that destination does with a caption track and by how likely the words are to change after you deliver.

Ship this
Because
What it costs you
Social cutthat autoplaysmuted
Burned in
Muted autoplay is the one thing a browser allows without a click,so the words have to be in the picture. Re-posts, downloadsand reshares carry the frames and drop the track.
The copy is frozen at render.Approve the words before youexport, not after.
Client reviewcut
Sidecar
The copy is still moving. A surname will be wrong, a job titlewill change, and somebody will want a line rephrasedafter the second viewing.
The client sees the player's default type,which reads as less finished thanthe burn-in they were picturing.
Broadcast orfestivaldeliverable
Whatever the spec sheet says
A delivery specification is not a preference. It names the format,and frequently the line count and the timing as well.Netflix's public style guide is a fair example of how specific they get.
Reading time, before the export.A rejected deliverable is asecond full pass.
A page onyour ownwebsite
Sidecar
You control the player, so you get a real caption track,a viewer toggle, and words that assistive technologycan read rather than pixels it cannot.
You have to keep the caption filewith the video for the lifeof the page.
A film going todistributors inseveral territories
A sidecar per language,never burned in
Burned-in English turns every other territory into a new master.Sidecars turn it into a new file against the same video.
Nothing worth naming.This one is not close.
Internalor archivecopy
Sidecar
The transcript is half the reason to keep the thing.A timecoded text file beside the master is worth morein two years than words baked into frames.
It only works if the two filesstay together, which is anaming discipline, not a feature.

Muted autoplay: Chrome's autoplay policy states that muted autoplay is always allowed, while autoplay with sound requires a prior interaction, a crossed engagement threshold, or an installed site. Delivery specifications vary by destination; the Netflix Timed Text Style Guide, cited here only as a published example of how prescriptive they get, sets 2 lines maximum per subtitle event, a minimum duration of five sixths of a second and a maximum of 7 seconds. Read the document you were actually sent.

The four expensive ones

Every mistake on this list
ends in a re-export and a second delivery.

Burning in before the copy is approved

This is the one that actually happens. The text is rendered, the film is delivered, and then somebody notices the founder's surname is spelled two ways. Now the fix is a render, an upload and an email to everyone holding the old file. Approve the words as text, in writing, before anything is burned. A caption file is the cheapest possible place to have that argument.

Burning in at one resolution, then downscaling

Type that is crisp at 3840 across goes soft at 1080, and softer again if the destination re-encodes what you sent. Thin weights and tight tracking suffer first. Burn at the resolution the cut will actually be watched at, or set the type heavy enough and large enough to survive the trip down, and check the small export on a phone rather than the timeline at full size.

A sidecar that gets separated from its video

The file arrives in a different email, or the client's upload ignores it, or the folder is unzipped and the small text file is left behind. Then the video ships with no captions at all, which is worse than either option on this page. Give the caption file the same name as the video, extension aside, keep both in one folder, and say in the delivery note that there are two files.

White text, no box, over a bright frame

Captions are read over whatever the shot is doing, and white type over a sky, a window or a snow exterior simply disappears. This is true of a burn-in and of a player's default styling alike. If you are burning in, add an outline, a shadow or a semi-opaque box and watch the whole film with it running. If you are shipping a sidecar, you cannot fix it at all, which is one more honest reason burn-ins exist.

What professional jobs actually ship

Both of them, from one source,
so the two versions can never disagree.

The answer on most real deliveries is not one format. It is burned-in on the cutdowns, sidecar on the master and the web version, and a single approved text behind both.

1

Approve the words once, as text

Transcribe the film, correct the names, the products and the jargon, and get a sign-off on the words themselves before any picture is rendered. This is the only step where a change is free. Everything downstream inherits whatever is in this file, including the mistakes.

2

Sidecar on the master and the web version

The long cut, the client's own copy and the page on their website all get the caption file. Those are the versions most likely to be revised, reused, translated, or opened by somebody who needs the words as words rather than as pixels.

3

Burn the cutdowns from the approved text

The vertical and square cuts get the type rendered in, because that is where the track gets stripped and where the type is doing design work rather than access work. Burn from the approved file, never from a fresh transcription pass, or the wording stops matching everywhere else.

4

Deliver the caption file even where you burned in

The client will need the words again: a re-cut, a translation, a transcript on their site, the next agency. It costs nothing to include and it is the item nobody asks for until the day it is missing and the only copy is inside a video.

The rule underneath all of it

One caption source, many renders. The moment the burned-in version and the sidecar are edited separately, they begin to drift, and the drift stays invisible until somebody watches two cuts back to back and asks which line is right.

Where this sits here

The caption track is text,
and the picture is never touched.

Captions here begin as a machine draft. A speech-to-text model transcribes the film and times the lines against it, and the track is stored and edited as WebVTT cues, which is to say as text rather than as pixels. You correct the names and the jargon in the editor, which is the pass no model does for you and the pass a client notices. Generation is metered in minutes per calendar month by plan, and the free Starter plan includes 20 of them, which is enough to hear what transcript quality looks like on your own audio before committing to anything. When the corrected track is saved, the caption panel hands you the file in both formats: a .vtt, which is how the track is stored, and a .srt, converted cue by cue on the way out. That is the sidecar this page has been arguing for, and it is the file you deliver alongside the video. One practical limit: a master parked in cold storage has no hot copy for the engine to read, so bring it back before you ask for captions. A second: a vaulted master refuses the caption download outright, because a subtitle file is a complete transcript of the dialogue.

Two things follow. The caption you make here is a sidecar by construction, so the film keeps exactly the picture you uploaded. Burning in is a job for your edit application, at export, before any of this happens. And a burned-in export is simply another file here, up to 5 TB per file subject to your available storage, streamed byte for byte with no re-encode anywhere in the path. That matters more for burned-in type than for anything else on a timeline: a platform’s second pass over your frames is what turns a clean caption into a soft one, and there is no second pass here. The mechanism is spelled out in video hosting that does not compress.

One claim this page will not make, because it did not survive checking: that a caption file will get the video found. Google’s own video documentation recommends structured data, video sitemaps, stable URLs, fetchable video files and key moments, and says nothing about transcripts or caption text. Ship captions for the people watching, and treat any ranking benefit as unproven until somebody publishes it. For the rest of the caption workflow, from the first generated draft to what lands in the client folder, start at captions for client video work.

Questions

Frequently asked

If I only get to pick one, which should it be?

A sidecar file, for almost everything that is client work. Client work gets revised, reused and handed to somebody else, and every one of those is cheap with a text file and expensive with pixels. Burn in when the video is going somewhere that plays it muted by default and strips caption tracks, which in practice means social cutdowns, or when the type is part of the design rather than an access feature.

Does burned-in text look better?

It can, and that is the honest argument for it. You choose the typeface, the weight, the size, the safe margins and the position, and you can animate it. A sidecar hands all of that to the player. SubRip carries no styling at all, and while the WebVTT specification defines cue settings for line, position, size and alignment along with STYLE blocks and ::cue styling from CSS, whether any of it is honoured depends on the player. If the type is doing creative work, burn it in and accept the permanence.

Can I do both on the same job?

That is what most professional deliveries actually look like: burned-in on the social cutdowns, a sidecar on the master and the web version. The rule that makes it safe is that both come from one approved caption source. Burn from the file that was signed off, never from a fresh pass, or the two versions quietly drift apart and nobody notices until a client watches them back to back.

What is the difference between .srt and .vtt?

SubRip (.srt) is a numbered list of timecodes and lines and nothing else, which is exactly why almost everything reads it. YouTube lists SubRip among the basic caption formats it accepts, alongside SubViewer, MPsub, LRC and Videotron Lambda. WebVTT (.vtt) is the W3C format browsers use natively; its specification defines cue positioning settings, STYLE blocks inside the file and styling through the ::cue pseudo-element, and it is served as text/vtt. Plan for SubRip's capabilities and treat WebVTT styling as a bonus where you control the player.

Are captions actually required, or just nice to have?

The Web Content Accessibility Guidelines put captions for prerecorded synchronised media at level A, the lowest conformance level there is: success criterion 1.2.2 Captions (Prerecorded). Live audio is 1.2.4 Captions (Live), at level AA. Whether an obligation applies to your client is a contract question, not a taste question, so ask which standard is named before you quote the job. Note that 1.4.5 Images of Text, at level AA, asks that text be used to convey information rather than images of text, with an exception when a particular presentation is essential. A burn-in is an image of text. Whether that is acceptable is for whoever signs off the audit, and this page will not pretend to answer it for them.

What happens if I burn in and the client then asks for French?

You render a second master, and a third for the next territory, and you keep them in sync forever. With sidecars it is one more small file against the same video, and the video is untouched. This is why a film heading to distributors in several territories should never carry burned-in dialogue: every language you bake in turns a file problem into a mastering problem.

Get the words right before anything is rendered

Caption generation is metered in minutes a calendar month by plan, and the free Starter plan includes 20 of them. Enough to transcribe a short piece, correct the names, and see what the text looks like before you decide what to burn.

Sources

  1. 1.W3C, Web Content Accessibility Guidelines 2.2, success criterion 1.2.2 Captions (Prerecorded), level A (fetched 9 September 2026)
  2. 2.W3C, Web Content Accessibility Guidelines 2.2, success criterion 1.4.5 Images of Text, level AA (fetched 9 September 2026)
  3. 3.W3C, Web Content Accessibility Guidelines 2.2, success criterion 1.2.4 Captions (Live), level AA (fetched 9 September 2026)
  4. 4.W3C, WebVTT: The Web Video Text Tracks Format, Candidate Recommendation Draft 20 May 2026 (cue settings, STYLE blocks, ::cue, text/vtt) (fetched 9 September 2026)
  5. 5.YouTube Help: supported subtitle and caption files, basic formats including SubRip (.srt) (fetched 9 September 2026)
  6. 6.Netflix Partner Help Center, Timed Text Style Guide General Requirements: 2 lines maximum per subtitle event, minimum duration five sixths of a second, maximum duration 7 seconds (fetched 9 September 2026)
  7. 7.Chrome for Developers, autoplay policy: muted autoplay is always allowed, sound requires prior interaction or engagement (fetched 9 September 2026)
  8. 8.Google Search Central, video SEO best practices: structured data, video sitemaps, stable URLs and key moments, with no transcript or caption recommendation (fetched 9 September 2026)
  9. 9.uncompressed.io: plans and caption allowances, and product behaviour read in the product source on 9 September 2026 (lib/facts.ts, app/api/captions/generate/route.ts, app/api/captions/[videoId]/route.ts, app/api/captions/[videoId]/download/route.ts for the .srt and .vtt download and the vault refusal, lib/vtt.ts, app/dashboard/Uploader.tsx)