How do you find every clip of one person in your footage?
You need face recognition, not a better description: visual search matches words against what a frame looks like, and no description survives a change of coat. Face recognition groups faces with each other, you put a name on each group once, and from then on the name finds the clips.
Published September 29, 2026 · Field notes
The question usually arrives in a specific form: every shot of the bride's grandmother across a two-day wedding, every clip of one interviewee across a three-year documentary, every appearance of a founder for a ten-year anniversary reel. The footage exists. The problem is that nothing in the filename, the folder or the picture description says who is in it.
Why describing the person does not work
Visual semantic search matches your words against the content of a frame. It is good at "woman in a red coat at a window" and useless at "Aunt Maria", because the model was never shown Aunt Maria and does not store identity. Even a detailed description breaks the moment she changes her coat. This is the one retrieval problem where describing harder does not help - you need a model that compares a face with other faces, not with text.
What face recognition actually does
It is two models, not one. A detector finds each face in a frame and the five landmarks around the eyes, nose and mouth. The face is then aligned and turned into a short vector by a recognition model, and faces whose vectors are close enough get treated as the same person. Nothing in that pipeline knows a name. Naming is always the human step: you look at a group, recognise who it is, and type it.
That also means there is no such thing as "search for a person you have never named". The first pass gives you anonymous groups; the value arrives after you spend a few minutes putting names on the ones that matter.
The four ways to do it today
- Scrub and mark by hand. Accurate, free and slow - roughly real time, which is fine for a single event and hopeless for an archive.
- Your NLE. DaVinci Resolve Studio offers people detection and face-based organisation inside a project's media pool, which is excellent if the footage is already imported into that project. See can DaVinci Resolve search your footage.
- A cloud service. Strong models, but every frame with a face in it leaves your machine. For client work, minors or anything under an NDA, that is often the end of the conversation - see does searching your footage upload it.
- A local library-wide index. Faces are grouped across every folder you index, not just one project, and the data never leaves the disk. This is where Reelary sits.
What Reelary does here, concretely
Reelary is our tool, so read this as disclosure. People recognition arrived in version 1.1 (September 2026) and is part of Pro. After indexing, a background pass re-extracts each indexed frame at 720p - the grid thumbnails are too small for faces - finds faces with the YuNet detector and embeds them with SFace, both open models from the OpenCV model zoo that run on the CPU. Faces are assigned to a person when they are close enough to that person's average, and the person list is kept on disk, so names and groupings survive re-indexing.
People appear as faces in the sidebar. Clicking one lists every clip that person appears in, as time ranges you can drop straight onto the timeline. Typing a name you have assigned into the search box returns that person's clips. Renaming a person to a name that already exists merges the two groups.
Where it stops
- Small faces are skipped. Faces under about 48 pixels across in the 720p frame are not recognised, because at that size the vector is mostly noise. Wide crowd shots and distant figures will not be attributed to anyone.
- It only sees sampled frames. Faces come from the same frames the visual index samples - shot boundaries or every 2 seconds, capped at 16 per video. A short clip is covered well; in a 90-minute recording someone who is on screen for ten seconds can be missed entirely.
- Profiles, masks and heavy grading. Side-on faces, sunglasses and strong colour casts lower the match rate, and the same person can end up split into two groups. Merge by giving both the same name.
- Lookalikes. Siblings and people of similar age and build can land in one group. There is no automatic fix; you will notice it when a person's clip list contains someone else.
When not to use Reelary for this.
If all the footage already sits in one Resolve Studio project, its built-in face organisation is right there and a second tool adds little. If you need face recognition that holds up on tiny faces in crowd shots, or across decades of ageing, larger cloud models are stronger, provided you are allowed to upload the material. And if the footage shows people who did not agree to being identified, check what your contracts and local law say about storing face data, even on your own disk.
Questions people ask about this
Can visual search find a specific person?
Not reliably. A frame-embedding model matches the overall content of a frame, so "woman in a red coat" works but "my sister" does not, because the model has never seen your sister and does not encode identity. Finding a specific person needs face recognition, which is a separate model that compares faces with each other rather than with text.
Does face recognition in Reelary upload faces anywhere?
No. Detection and recognition run on your machine, the face data is stored in the local index next to the rest of your library, and nothing is sent to a server. You can index and name people with the network disconnected.
What if it splits one person into two, or puts two people together?
Both happen. The same person can land in two groups when lighting, age or angle differ a lot, and people who look alike can share a group. Renaming one group to the name of another merges them, so the usual fix for a split is to give both the same name.
Is people recognition free?
No, it is part of Reelary Pro, along with dialogue search. Visual search, AI scene descriptions and unlimited indexing are free.