Captions and search
Search for a word inside a video
and land on the frame where it is said.
Scrubbing is how most people look for a line, and it is the slowest tool in the building. A caption track changes the job: the dialogue becomes text with timings attached, so a word you half remember becomes a query, and every place it was spoken becomes a button. Here is what that search does, what it saves on real work, and the two things it cannot do.
Video hosting for filmmakers · 5 GB free · Paid plans from USD 9/month
Updated September 2026
What the search actually is
A caption track is
an index of everything anybody said.
Captions are usually filed under accessibility and delivery, which undersells them. A caption track is the spoken content of a film written down, with each line pinned to the moment it was said. Once that exists, the film stops being an opaque hour of pictures and becomes something you can query.
The mechanism is deliberately small. The caption panel carries a search field. What you type is matched against the caption cues, ignoring case and accents, so a search typed without accents still finds an accented word and a search in lower case still finds a name written in capitals. Every matching word is rendered as a button, and pressing it seeks the player to that moment.
That is the whole feature. It is worth being precise about how little it is, because the value is not in the sophistication of the search. The value is that the question “where does she say the thing about the warehouse” used to be answered by a person dragging a playhead, and is now answered by typing a word.
1
Field, in the caption panel
It sits with the captions rather than in a separate transcript tool
0
Accents you have to type
Matching ignores case and accents in both directions
1 click
From a match to the frame
Every matching word is a button that seeks the player
20 min
Caption minutes, free plan
Generation is metered in minutes a month by plan
Getting there
From an uncaptioned film
to a searchable one.
Four steps, and only one of them takes any of your attention. The third step is the one people skip, and it is the one that decides how much the search is worth to you later.
Upload the cut
The film goes up as the export you made. Nothing is re-encoded on the way in, so what a reviewer watches is the file you approved. Captions are generated against that film, so caption once per cut you actually intend people to watch rather than once per rough pass.
Generate the captions
A speech-to-text model writes the cues against the film's own timeline. Generation is metered in minutes per month by plan, and the free Starter plan includes 20 minutes a month, so a short piece or a few interview selects cost you nothing to try.
Read them once and fix the proper nouns
This is the pass that matters. Names, company names, product names and trade jargon are exactly what people search for and exactly what a speech model gets wrong. The owner edits the captions directly, and a reviewer who spots an error can submit a correction that the creator applies, which is often faster than the editor catching it alone.
Search from the caption panel
Type a word into the field in the caption panel. Matching words appear as buttons through the transcript. Press one and the player seeks to that moment. Nobody needs to know a timecode, and nobody needs to ask the editor for one.
Captions are the prerequisite, not the subject of this page
What it replaces
Four jobs that cost hours,
a search box costs seconds.
None of these are hypothetical. They are the four reasons people actually type into a caption panel, and in each one the manual alternative is a person watching a film with a notepad, at the speed the film plays.
Times are arithmetic on your own footage rather than a measured benchmark: the manual column is the runtime you have to sit through, the search column is a query. Every row assumes the word you type was transcribed correctly, which is why the editing pass in step three is not optional.
How to search well
Type the word you know was said,
not the sentence you remember.
The search matches text in the caption cues. That is a literal thing, not a clever one, and knowing it makes the difference between finding a line in one try and concluding the feature does not work.
One distinctive word beats a sentence
Cues are short lines. A sentence you remember may be split across two of them, and your memory of the wording is probably not exact anyway. Pick the least common word in the line, the one nobody else in the film would say, and search that.
Search the words, not the idea
It matches what was spoken, not what was meant. A search for pricing will not find a passage where somebody talked about pricing without using the word. If you are sweeping for a topic rather than a phrase, search several of its likely words in turn.
Accents and case are free
You never have to reproduce the diacritics or the capitalisation of the transcript. Type it flat and fast. On French dialogue reviewed on an English keyboard this is the difference between using the search and giving up on it.
Read this before you rely on it
The search is only as good
as the transcript underneath it.
A word the model misheard is a word the search cannot find. If an interviewee says a company name and the transcript writes something that merely sounds like it, that mention does not exist as far as the search is concerned, and nothing in the interface will tell you it is missing. This is the honest limit of every dialogue search built on machine transcription, ours included, and it is the precise reason the editing pass on the captions is worth the twenty minutes it takes.
The practical rule follows from that. Do not treat a search that returns nothing as proof that a phrase was never said. For a casual lookup, an empty result is usually correct and always cheap. For a legal or compliance question, the search is a way to find the mentions fast, not a certificate that there are none, and the certificate still comes from a person who read the transcript. Correct the names and the jargon first, then the two uses converge.
The second limit is about what is being searched at all. This reads spoken dialogue as transcribed into the caption cues. It does not read text burned into the picture, a lower third, a title card, a slate or anything in a graphic. It does not search the image: there is no way to look for a face, an object, a location or a shot type, and we are not going to imply otherwise. If your problem is finding a shot rather than a line, this feature is not the answer and your bin structure still is.
Where this sits in the workflow
Questions
Frequently asked
Can I search a film that has no captions?
No. The search reads the caption cues, so a film without a caption track has nothing to match against. Generate captions on the film first, then the search field appears with them in the caption panel.
Do I have to type the accents correctly?
No. Matching ignores case and accents, so typing deja finds a cue that reads deja vu with its accents in place, and typing MONTREAL finds it written any way it was written. This matters most on French and Spanish dialogue, where the person searching is often on a keyboard that makes accents awkward.
What happens when I click a result?
Every matching word in the caption panel is a button. Pressing it seeks the player to that moment, so you go from a query to the frame in one click rather than reading a timecode off a list and typing it into a player.
Does it find words that appear on screen, like a lower third or a slate?
No. It searches spoken dialogue as transcribed into the caption cues. Text burned into the picture, text in a graphic, a name on a slate and anything visual are all invisible to it. There is no visual search here: you cannot search for a face, an object, a location or a shot type.
What if the model misheard the word I am looking for?
Then the search cannot find it, and no amount of retyping will change that. The search is only as good as the transcript underneath it. This is the practical reason to read the captions once and correct the proper nouns, the product names and the jargon before anybody relies on the search. The owner edits the captions, and a reviewer can submit a correction that the creator applies.
Can the client search, or only me?
Whoever can open the film can use it. A review room is opened with the share link, which is itself the authentication, so a client or a producer searches the dialogue without an account, a password or a seat. That is the point of the feature for most studios: the person who remembers the phrase can find it themselves.
How many minutes of captions do I get?
Caption generation is metered in minutes per month and the allowance depends on the plan. The free Starter plan includes 20 minutes a month, which is enough to caption a short film or a couple of interview selects and judge whether the search earns its place in your workflow.
Caption a film once, then stop scrubbing for lines
Caption generation is metered in minutes a month by plan, and the free Starter plan includes 20 minutes a month. The search in the caption panel comes with them, on every plan.
