How do you search what was said across all your footage?
Transcribe every file in the library, not just the project you have open, and search the transcripts two ways: by exact words when you remember the quote, and by meaning when you only remember the point. A multilingual model makes the second one work across languages too.
Published September 29, 2026 · Field notes
Anyone with interviews in their library has had this moment: you remember roughly what someone said, maybe the point they made rather than the words, and you have forty hours of rushes. The picture will not help - every frame is the same person in the same chair. What you need to search is the audio, and it has to cover the whole library, not just the project you have open.
Why a transcript beats a picture here
Speech is dense in a way images are not. One sentence carries names, numbers and a claim; a frame of a talking head carries a talking head. That is why, for anything someone said, the transcript is the index to build first - the fuller argument is in the difference between searching your footage and transcribing it.
The catch has always been the workflow. Transcribing one interview is easy. Transcribing every file on three drives, keeping it up to date as new shoots arrive, and searching it all from one box is the part most setups never get to.
Exact words versus meaning
Keyword search over a transcript is precise and trustworthy: the word is there or it is not. Its weakness is that you rarely remember the exact words. You remember "the part where she talks about quitting her job", and she actually said "the day I walked out of the office". A keyword search for "quitting" finds nothing.
Searching by meaning fixes that. A text embedding model turns each transcript segment and your query into vectors, and segments that mean something close to the query come back even when they share no words. The same trick crosses languages: a multilingual model places "the festival at night" and a Spanish line about the fiesta near each other. The price is that meaning is fuzzy. Use it when you remember the idea, and exact words when you remember the quote.
What Reelary does here, concretely
Reelary is our tool, so read this as disclosure. Dialogue search arrived in version 1.1 (September 2026) and is part of Pro. After a one-time download of the speech model, Reelary transcribes every indexed video in the background with whisper.cpp running on your machine. It covers the whole audio track, not the sampled frames the visual index uses, so a 90-minute recording is searchable end to end. The language is detected per file. Files are re-transcribed only when they change, and silent files are marked so they are not retried.
Transcripts are searched two ways at once: by the words you typed, with a bonus for the exact phrase, and by meaning, using the multilingual-e5-small model, also on device. Meaning-only matches only count when they stand clearly above the rest of the library for that query, which keeps "car" from returning every line that mentions anything with wheels. Results arrive as time ranges in the same grid as visual results, so a line of dialogue can go straight onto the timeline.
The same text search also covers AI scene descriptions, which are free: a local vision-language model writes a one-line description of each indexed shot. That means a query like "priest blessing the crowd with water" can match a description even when nobody in the shot spoke.
Where it stops
- Recognition errors. The speech model is a compact one chosen to run on ordinary laptops. It handles clear speech well and mishears names, jargon and noisy audio. A misheard word will not match exactly; search by meaning is the fallback.
- No speaker labels. It knows what was said and when, not who said it.
- Time to catch up. The first pass over a large library runs in the background and takes a while on older machines. New files are picked up as they are indexed.
- Not a transcript editor. You search the transcript; you do not edit it, export it or cut by deleting text.
When not to use Reelary for this.
If your work is almost entirely interviews and podcasts and you want to edit by deleting words from a transcript, with speaker labels and exportable text, a text-based editing tool is the better fit. If the footage is already in a Premiere or Resolve project, their built-in transcript search is right there. Reelary is for searching what was said across a whole library, alongside the picture, before you know which project it belongs to.
Questions people ask about this
Do I need to know which language the footage is in?
No. Reelary transcribes with a multilingual Whisper model that detects the language of each file by itself, so a library that mixes English interviews with Spanish or German street footage needs no setup.
Can I search in English and find a line spoken in another language?
Often, yes. Transcripts are indexed twice: by exact words, and by meaning with a multilingual text embedding model. The meaning index is what lets "village festival at night" find a Spanish line about a fiesta. An exact phrase match still ranks above a meaning-only match, so a quote you remember word for word comes first.
How accurate is the transcript?
Good on clear speech from a lav or a close mic, noticeably weaker on crowd noise, music, heavy accents and proper nouns. A name the model misheard will not match an exact search, though searching by meaning can still find the line.
Does it label who is speaking?
No. There is no speaker diarisation, so the transcript does not know who said a line. To find a line by one person, combine it with people recognition, or use a dedicated transcription tool.