Glossary

Glossary

The words behind transcription and media search, explained in plain language with examples.

๐ŸŽต
Audio fingerprinting

What is audio fingerprinting?

How song-recognition apps identify a track from a few noisy seconds, and why a cover version slips past them.

Read the definition โ†’
๐ŸŽฅ
B-roll

What is B-roll?

The supporting shots that cover cuts and show what's being talked about, and the real problem: finding them later.

Read the definition โ†’
๐Ÿ”ก
Captions vs subtitles

Captions vs subtitles: what's the difference?

Why captions describe sounds and subtitles don't, plus open vs closed captions and SDH, with examples of each.

Read the definition โ†’
๐ŸŽž๏ธ
EDL (edit decision list)

What is an EDL (edit decision list)?

The decades-old plain-text cut list every video editor can still read, and what it leaves out.

Read the definition โ†’
๐Ÿงฎ
Embeddings

What are embeddings?

The numeric fingerprints behind semantic search, explained without the math.

Read the definition โ†’
๐Ÿ‘ค
Face recognition

What is face recognition?

How software finds a face and matches it to the same person again, where it gets people wrong, and why consent matters.

Read the definition โ†’
๐ŸŽฌ
FCPXML

What is FCPXML?

Final Cut Pro's timeline interchange format, what it carries, and how it compares to an EDL.

Read the definition โ†’
๐Ÿ”€
Hybrid search

What is hybrid search?

Running exact keyword search and meaning-based search together, then fusing the results so neither blind spot wins.

Read the definition โ†’
โš–๏ธ
Keyword vs semantic search

Keyword vs semantic search

Exact-word matching versus meaning matching: when each one wins, and why the best results usually come from both.

Read the definition โ†’
๐Ÿฆ™
Local LLM

What is a local LLM?

A language model that runs on your own machine: what hardware it needs, what it does well, and where cloud models still win.

Read the definition โ†’
๐Ÿ”–
Logo detection

What is logo detection?

Spotting brand marks in photos and video, whether you search by a brand's name or by a sample of the logo.

Read the definition โ†’
๐Ÿ”Œ
MCP (Model Context Protocol)

What is MCP (Model Context Protocol)?

The open standard that lets AI assistants call tools on your computer, and what that means for your data.

Read the definition โ†’
๐Ÿ—„๏ธ
Media asset management (MAM)

What is media asset management (MAM)?

What MAM systems do, how they differ from DAM, and when a lighter local search tool is enough.

Read the definition โ†’
๐Ÿท๏ธ
Named entity recognition

What is named entity recognition (NER)?

Pulling the people, places and organizations out of text, and what changes when that text is a transcript.

Read the definition โ†’
๐Ÿ‘ฏ
Near-duplicate detection

What is near-duplicate detection?

Finding photos and videos that look the same even when the files differ, and how to clean them up safely.

Read the definition โ†’
๐Ÿ”ค
OCR (optical character recognition)

What is OCR (optical character recognition)?

How software reads text inside photos, screenshots and video frames, and what it takes to make that text searchable.

Read the definition โ†’
๐Ÿ’ป
On-device AI

What is on-device AI?

Running AI models on your own hardware instead of a server, and the real trade-offs that come with it.

Read the definition โ†’
๐Ÿชถ
Proxy media

What is proxy media?

Light stand-in copies of heavy camera files, how editors swap them back for the originals, and what breaks the swap.

Read the definition โ†’
๐Ÿฅ‡
Reranking

What is reranking?

The careful second pass that reorders a search's top candidates, and why it can polish results but never find missing ones.

Read the definition โ†’
๐Ÿ“š
Retrieval-augmented generation (RAG)

What is retrieval-augmented generation (RAG)?

Search first, then answer: how AI assistants answer questions about your own files, and why citations matter.

Read the definition โ†’
๐ŸŽฌ
Scene detection

What is scene detection?

How software spots cuts and transitions in video, what fools it, and why segmenting footage makes it searchable.

Read the definition โ†’
๐Ÿง 
Semantic search

What is semantic search?

Searching by meaning instead of matching words, and why it helps most when you can't remember exactly what was said.

Read the definition โ†’
๐Ÿ—ฃ๏ธ
Speaker diarization

What is speaker diarization?

How software works out who spoke when in a recording, and why the labels are anonymous until you name them.

Read the definition โ†’
๐Ÿ’ฌ
SRT vs VTT

SRT vs VTT: what's the difference?

Two plain-text subtitle formats that look almost identical, the small differences that break players, and which to pick.

Read the definition โ†’
๐Ÿ•
Timecode

What is timecode?

The HH:MM:SS:FF address stamped on every frame, why drop-frame exists, and how it differs from a player's timestamp.

Read the definition โ†’
๐Ÿ“
Transcription

What is transcription?

Turning speech into text, how automatic transcription works, and what makes a transcript actually useful.

Read the definition โ†’
๐Ÿ“ฆ
Vector database

What is a vector database?

The storage layer that makes semantic search fast: how it finds the nearest embeddings, and the trade-offs of doing it approximately.

Read the definition โ†’
๐Ÿ“‘
Video chaptering

What is video chaptering?

Breaking a long recording into titled sections, how software finds the breaks automatically, and where it misplaces them.

Read the definition โ†’
๐ŸŽฏ
Word error rate

What is word error rate (WER)?

The standard score for transcription accuracy, how to calculate it by hand, and why a single number hides what matters.

Read the definition โ†’
โฑ๏ธ
Word-level timestamps

What are word-level timestamps?

Timing for every single word in a transcript, and why that precision matters for search, captions and editing.

Read the definition โ†’

More: Use cases ยท Guides ยท Alternatives ยท Free tools