Changelog

What's new in MediaFind

Every release of the private, on-device media-search app — the features, changes, and fixes in each version. Builds are published on GitHub Releases.

Changed
  • Search answers a little faster. Each search now asks the image model once instead of three times, and reads only the matching part of a long summary when ranking results.
Fixed
  • The People count no longer includes people with no faces. Re-indexing a file used to leave empty entries behind that inflated the count and slowed the People page.
  • Searching while files are being added no longer pauses for seconds. The search index is rebuilt in the background after each new file instead of inside your search.
  • Faces & People opens on a large library. Its voice-to-face suggestions used to compare every voice with every person and could keep the page waiting for hours.
  • MediaFind says when something it needs is missing, instead of quietly doing less — for example when it cannot tell similar voices apart.
  • A problem report explains what happened before the restart. It now carries the app's recent log and a check of what it can and cannot do on your machine.
  • A step that fails while the rest of a file is added is written down, so a report can say what went missing instead of looking like nothing happened.
0.1.57
Added
  • A visual summary of any video or recording, when you ask for one. Choose Create visual summary on a file for a page of charts, takeaways and key numbers, each linked to its moment.
  • Say when a later meeting changed a decision, and Ask tells you. Mark a related decision Same decision, changed or Not related; when Ask quotes a decision that changed, it says so.
  • A notes folder you import now stays in sync while folder watching is on: notes you add or change are read again, and notes you delete are forgotten.
  • A Feedback button in the top bar of Find, Meetings and Memory, and on a search that finds nothing. Your note opens in your mail app, and nothing is sent until you send it.
  • Ask answers each question in a message that asks two or three, each under its own heading, as if you had asked it alone.
Changed
  • Suggested searches load faster on a large library. On a test library with a million name mentions, they took about 0.8 s on every page load and now about 13 ms.
  • MediaFind opens about three seconds sooner when your library holds files. The search model now loads after the window appears, in time for your first search.
  • You can ask your next question while Ask is still answering. It waits as Queued and goes out on its own once the answer is in.
Fixed
  • MediaFind opens again on the newest macOS. It quit the moment its window finished loading, on macOS updates released after 0.1.56.
  • Call on this Mac records your voice too. In the Mac app it recorded only the other people on a call, because the app shipped without the part that opens the microphone.
  • Re-transcribing, re-indexing and re-detecting speakers say so when that work is already paused. They used to queue it again, and resuming then did the work twice; now they point you to Resume in the activity dock.
  • Using the on-device AI just as it unloads after ten idle minutes no longer crashes MediaFind; it loads again and answers.
  • Searches understand a list of months, and a month with its year. "Videos in March and May", "from March" and "in March of last year" now mean what they say.
  • Each question you ask moves to the top of the Ask page, with its answer below it, instead of hiding under a pinned header.
  • A follow-up that reads like a whole question still follows the conversation: after an answer about Jupiter, "when is the best time to look?" stays about Jupiter.
  • Find's date filters find a note on the day its card shows.
  • Messages no longer say that only the Mac app can pick a folder, read on-screen text or offer phone access. The Windows and Linux apps do all three too.
  • A keynote, a match or a film no longer turns into a meeting you didn't ask for. Detect speakers now only finds who speaks, and Not a meeting takes a recording out of Meetings.
  • A video you download is dated when it was published, and Meetings leaves it out. Its timeline date and due dates follow the publish date, and Process recordings skips it unless you pick it.
  • A "redact PII" export of your conversations no longer gives away what it hides. Note titles taken from file names are withheld, and every redacted export masks more phone number and email formats.
  • Indexing on another Mac speeds things up again, and is never why a file fails. If that Mac can't take a file, this one indexes it, and your access token and licence stay here.
  • Quitting while the on-device AI is answering is clean again, and an answer or summary cut short by quitting is no longer saved as if it were complete.
  • A saved XL choice loads the default AI model on a Mac with less than 32 GB of memory, so a library moved from a bigger Mac keeps working.
  • The Summary backend setting now applies to files you add or download, not only to summaries filled in later.
  • A paused job stays paused and stays in the activity dock, even across a restart, and indexing a folder whose job is paused joins that job instead of queuing it twice.
  • Cancelling Detect objects stops it, and it no longer wipes boxes saved while it ran. Indexing also names any file it indexed without logo and action detection.
  • Ask no longer quotes a conversation you have forgotten, even one forgotten from the command line or by an AI assistant, and leaves out conversations that don't match the question.
  • Memory reads facts about you more carefully. "Please call me ASAP" no longer becomes your name, and each fact in Settings → Memory shows its conversation's title.
  • Library repairs are more reliable. Re-scan speakers now covers older recordings too, and moving a file in or out of quarantine on a full disk leaves no half-copied file behind.
  • Downloads refuse internal network addresses however they are written.
  • The desktop app's "downloader is out of date" message says to update MediaFind, instead of naming a developer command.
  • Cleanup → Storage counts conversation libraries and backups too; leaving them out could drop gigabytes from the total.
  • "Fill in the required fields." stands out again in form dialogs, instead of looking like the hint under a field.
  • Removing a file or deleting picked faces deletes their images, even after the data folder has moved. Images another file still uses are kept.
  • A file you move and relink keeps its own thumbnails, and removing a file never deletes another file's pictures. A file already showing another's pictures gets its own back when you re-index it.
0.1.56
Added
  • The status the app and the phone companion poll says whether folders are watched. GET /api/status (and its ?light=1 form) now carries watching, the folder watcher's {enabled, folders, last_scan_at}, as GET /api/settings already did. The Python SDK's Engine.search hits now say what kind of file they come from (kind: video, audio, image or text) and a conversation's hit its title, as /api/search hits do.
  • MediaFind keeps watching the folders you add. Only the command line (mediafind watch) could follow a folder before. Now, while the app runs — the desktop app or mediafind serve — it checks the folders in your library about once a minute and indexes new and changed files dropped into them, as ordinary index runs, so the free plan's file cap and demo mode apply as they do to Add Media. A file waits until it has stopped changing between two checks, so a copy or a recording still being written is left alone. Files added while the app was closed are picked up at its next start, a file you removed from the library stays removed, and a renamed file keeps its transcript instead of being transcribed again. On by default: the watch_folders setting, or MEDIAFIND_WATCH_FOLDERS=0, turns it off without a restart.
  • The memory API's responses are pinned and published. Every /api/v1/memory route now answers with a declared schema, listed in the server's /openapi.json, so an assistant or evaluation harness built on it can rely on each field: within v1 fields are only added, never removed, renamed or retyped. A field a route can leave out is still left out rather than sent empty. The responses themselves do not change; they are now held to what they already were.
  • A memory question about one month can now be held to that month. Asked "what did I decide in April?", memory search took April as the point to start from and let every message since then compete for the answer — right for a chat where you talk about April in June, wrong for notes, photos and messages that carry the date they belong to. The end of the period can now be applied too: --question-span span on mediafind memory search and mediafind memory ask, question_span in the memory ask endpoint (the search endpoint already had it), and the same option on the memory_search tool an assistant uses over MCP. It also reads dates the start-only reading had to ignore, such as "early March" and "March 14". Every surface still defaults to what it did before.
  • Choose how facts are read from your conversations, in Settings → Memory. What people say about themselves — where they live, what they bought — is read from each saved conversation with built-in patterns, with the on-device model, or not at all. That choice used to be an environment variable read once at start; Read facts with now sets it, and it applies to the next conversation you save, without a restart. An environment variable still overrides it, and the setting says so.
  • A meeting's decisions show related decisions from your other meetings. Under each decision in a meeting's brief, MediaFind lists decisions from earlier and later meetings about something close — "We decided to move the beta launch to September" beside May's "We decided to launch the beta in May" — each dated, and a click opens that meeting where it was said. They are shown as related, not as replacements: whether the later one changed the earlier is yours to judge. Keyless and on-device, for the meetings you have processed.
  • The command line catches up with the app. mediafind status prints your plan, how many files are indexed (against the free plan's cap), the version and where the library lives — the fields of GET /api/v1/status, as JSON with --json. mediafind license shows your plan and subscription, mediafind license activate takes the key from checkout (typed after the command, or read from stdin so it stays out of your shell history) and mediafind license deactivate removes it (without cancelling the subscription): the same steps as the app's Settings. mediafind search now prints the filters it read from your query ("Interpreted as Video — searched “bread”"), searches the words as typed with --no-parse (or --literal), and takes --after, --before and --sort like the API; when the filters it read leave nothing, it says so and points at --no-parse instead of asking whether you indexed a folder. mediafind export gains the app's speakers and clean_transcript formats.
  • The Mac app reads on-screen text. Slides, captions, signs and other text in your videos and images become searchable, and the on-screen text download has something to download. The app now includes an open-source OCR engine and its English model instead of relying on a copy you installed yourself, which the app could not find anyway. The model is English: text in other alphabets, such as Chinese, Arabic or Cyrillic, isn't read yet.
  • Pair the phone app by scanning a QR code, and it tells you what to fix when it can't connect. Scan the code in MediaFind → Settings → Phone access (or open its mediafind://pair link) and the phone saves the address and access token together, once the computer accepts them. The app used to look on port 8000 and tell people to set an environment variable; it now uses phone access's port, 7860, and its Wi-Fi scan tries 7860 and then 8000 and shows each computer's MediaFind version. A refused token says whether phone access is off or the token changed, and Settings → Re-pair / change token takes a new token without forgetting the address. Conversations found by search open as text rather than as a file that can't play, files play on the phone when MediaFind on the computer offers playback links, Copy path is a real button, and the Both filter, which ran the same search as All, is gone. These need version 1.1.0 of the phone app.
Changed
  • The website and README present MediaFind as four parts: Find, Meetings, Memory and Create. The homepage keeps finding a moment as the way in, then gives each part its own section, all reading one local library: Meetings for live and recorded meeting notes, Memory for ChatGPT and Claude exports, pasted chats and folders of notes you can search, ask and delete, and Create for the timeline editor. It now says photos are indexed too, and every page's Features menu lists the four parts.
  • One list of what Pro adds. The pricing section, the feature cards and the README each named a different set. Now all three name the same features, by the names the app's paywall uses, including the ones they left out (collections, the knowledge map, highlight reels, brand logos, the downloader and Create's assembly tools), and the pricing section and README add unlimited indexing. make sync fails when a list names anything the app doesn't gate or drops one of them. The free plan lists only free features.
  • PhotoFind, MediaDownload, MediaClean and MediaTranscribe are free. The menu that lists them is now *Free apps*, and each app's page says it is free and what MediaFind adds. PhotoFind's page no longer sells a Pro upgrade: its checkout gave buyers no licence key. MediaDownload's page no longer promises downloads from sites such as YouTube, which its current build may fail on.
  • The download section says what works where. Windows and Linux are labelled beta, and it says they include ffmpeg, an OCR engine and phone access, as the Mac app does, and that the Mac app needs macOS 14 or later. The phone companion is described as what the current Android build does: search and read transcripts from MediaFind on your computer.
  • Pair your phone by scanning a QR code. While phone access is on, Settings → Phone access (the card was "Access from your phone") shows a QR code the MediaFind mobile app scans: it carries this computer's address and the access token together, so nothing has to be typed. It is drawn on this computer, with nothing fetched. Copy pairing link copies the same link, and where this computer has more than one network address you pick the one your phone's Wi-Fi reaches.
  • Settings says how to get moments into Premiere Pro. There is no Premiere panel to install from the app yet, so Editor integration says what works today: Export to editor → FCPXML in Find, then *File ▸ Import* in Premiere.
  • Settings → Preferences has a "Watch my folders for new files" switch for the watch_folders setting, with a line saying how many folders MediaFind watches and when it last looked. It shows where the settings offer folder watching, and says so on hover when MEDIAFIND_WATCH_FOLDERS pins it.
Fixed
  • The Windows and Linux builds read on-screen text. Finding and downloading the text on slides, signs and burned-in subtitles runs an OCR engine, which those builds never included: unless it happened to be installed and on the PATH, they read no on-screen text and said nothing. They now carry the engine and its English model — on Linux a self-contained build that needs glibc 2.35 or newer — and a build without them fails instead of shipping. Only English text is read so far.
  • The Windows and Linux apps work like the Mac app. Their installers shipped without ffmpeg, so Create renders, the Video Editor and clip export failed; they now include it, and the build fails without it. The app keeps one address (port 7860 when it is free, else the one it used last), so your theme, dismissed notices, recent searches and Ask conversation survive a restart; "Choose folder…" and every "save to a folder" open a real folder picker; Access from your phone works; a second launch brings the open window forward instead of starting a second engine; and quitting lets the engine shut down, instead of killing it and leaving several GB in the temp folder. On Windows no console window opens beside the app, and when the engine doesn't start the app names the likely cause (a library made by a newer version among them) and where its logs are. The server microphone is included, and an NVIDIA machine whose CUDA can't run the transcription model now transcribes on the CPU instead of failing every file.
  • mediafind ask says when it withheld an answer. An answer the sources don't support printed "No on-device model — open Settings to choose a model" above "The closest moments don't answer that.", then the closest sources as its "Sources". It now prints the headline and why, as the Ask page does, and lists no sources for an answer it did not give (with --clips, it names the clips of the closest moments it cut). An answer is labelled grounded or not grounded as the answer itself says, instead of every model answer "grounded" and no quote ever "not grounded", and an explanatory answer ("… is not in the index") no longer claims a model is missing. --after/--before say they take an ISO-8601 date-time too.
  • A saved search with a Pro channel still refreshes on the free plan. A search saved during the trial keeps its People, Brand logos or Shadow channel (the Find page saves a Pro user's recommended channels whole), and refreshing it answered 402 once the trial ended. It now runs with its other channels and says which it left out; only a saved search of Pro channels alone, or one with a person filter, is still refused.
  • The assistant tools and the SDK match the other surfaces. The ask_library tool retrieved 5 sources where every other surface retrieves 6; search_media answered an after or before it could not read with a tool error, and now returns an invalid result saying why; library_status reports the MediaFind version. The MCP guide and the SDK guide describe withheld answers (abstained, reason), the Pro search channels, the date formats, mediafind license, and the refusal of a library a newer MediaFind made.
  • The Generate room shows when local still generation is on. Create hid the room unless its page was marked for generation, and nothing marked it, so MEDIAFIND_GEN_SLICE=1 left the room out of reach.
  • "On this day" counts files, and Memory's errors say space. The Home rail's "On this day" card said "3 memories from 2 past years" about files, and Memory's messages said "library" where the Memory page says "space" ("another conversation import is already running for this space").
  • Content Credentials are "unavailable", not "absent", when the inspector cannot start. A c2patool whose loader could not find a library it links printed "… not found", which was read as a file with no credentials. It is now reported as unavailable, with the loader's reason.
  • The website says MediaFind watches your folders, as it now does. While the app is open it checks the folders you added about once a minute and indexes new files once they have finished copying; Settings turns it off.
  • The browser extension, the Premiere panel and the bookmarklet find MediaFind. They looked for it on port 8000, where mediafind serve listens, while the desktop app serves on 7860 (or a free port when 7860 is taken). A running MediaFind now records where it listens in ~/.mediafind/server.json, and GET /health says it is MediaFind and which client versions it serves. The Premiere panel reads that file and remembers the address; the extension asks 7860, then 8000, then an address you enter; the bookmarklet the app hands out targets the port that app listens on. The extension no longer says "Captured" when the download has only been queued: it follows it until it is indexed or fails. The Premiere panel no longer offers to put a conversation on the timeline, because search results now say what kind of file they come from. A phone that is refused now learns whether its token is wrong or phone access is off on the computer, and GET /api/v1/media-link gives the companion app a signed link to play a file from, which works for an hour without the token header its video player cannot send. The DaVinci Resolve script stops with a message to reinstall it when a newer MediaFind wrote the handoff, instead of misreading it.
  • Ask and search answer the same way on every surface. Asked something your conversations don't cover — "How much did the new car cost?" over messages about a cat, a kitchen and a report — the Ask page said "Your conversations don't cover that.", while the API (/api/v1/ask), GET /api/ask, the Python SDK, the ask_library tool an assistant uses over MCP and mediafind ask quoted the cat message as a grounded answer. They now withhold that answer as the page does, and their answers carry abstained and reason. grounded is true only for an answer its sources support, not for the nearest line quoted for want of one, and backend: "extractive" never runs the local model. Measured on the Ask fixtures, all 92 answerable questions still answer, and 14 of the 16 unanswerable ones are no longer reported as grounded. On /api/v1/search, speaker and person now apply. A misspelt field is named in a new warnings list instead of being ignored silently, and an unknown modality or a blank query is refused. after and before, there and on /api/search, take ISO-8601 date-times and read the date a file was captured (else modified), as a date typed into the query does; one that doesn't read is refused instead of dropped. The Shadow channel is Pro in search, as it is everywhere else, and the SDK's Engine.index takes only its documented options and refuses meeting=True on the free plan, as mediafind index --meeting does.
  • An older MediaFind can no longer change a library a newer one made. The app has always refused to open such a library; the command line, the MCP server an AI assistant starts, and the Python SDK opened it anyway and applied their own, older schema changes to it — which is what happened when an assistant's MCP config ran an old source checkout against the desktop app's library. They now stop with the app's message, which says how to get back; mediafind restore and mediafind backup still run.
  • MediaFind installed into a Python environment uses the desktop app's library. Installed from a wheel rather than run from a source checkout, it kept its data inside that environment, where it saw neither the app's library nor its licence, and lost both when the environment was rebuilt. It now uses the per-user data directory the app uses (~/Library/Application Support/MediaFind on macOS); a source checkout, editable installs included, still keeps data/, and MEDIAFIND_DATA still overrides both. A library such an install indexed before stays in the environment's site-packages/data: set MEDIAFIND_DATA to that folder to keep using it.
  • Transcripts no longer carry subtitle credits nobody said. Over music or silence the recognizer sometimes writes the things it learned from subtitle files rather than from speech — "Sous-titres par …", a broadcaster's © stamp, a streaming platform's watermark, a transcriber's byline, or a "thanks for watching" stretched across half a minute — and because those are distinctive strings, searching for one found it. They are now dropped as a file is transcribed, and from a live meeting's transcript (a call's hold music is where they turn up). Only lines that are the credit itself go: a real sentence that happens to end with a watermark keeps its words, someone talking about subtitles or reading out a copyright notice is left alone, and so is a presenter actually saying "thanks for watching". Measured on a 13,163-segment library, it removes 13 invented lines and touches nothing that was said. Existing transcripts are repaired by re-transcribing the file; MEDIAFIND_ASR_CREDIT_FILTER=0 turns it off.
  • A memory question that names two months finds what happened in the first one. Asked "what did I decide in March and in May?", memory search and Ask took the later month alone and searched from May on, so everything said in March was left out — the very part the question asked about. A question naming several periods now searches all of them, and one period inside another ("in 2023, in March") still means the inner one.
  • Find shows the files from both months when you name two. Searching "videos in March and in May" treated the two months as conditions that must both hold, and since nothing was taken in both, Find showed nothing at all. A search that names several periods now finds files from any of them, and one period inside another ("in 2023, in March") still means the inner one.
  • The app says when it can't read on-screen text. Reading text off the screen needs an OCR engine, yet Add Media called it "on by default" wherever the engine was missing. Where it can't be found, Settings and Add Media now say so beside the OCR switch, and mediafind doctor lists it.
  • The website stops promising what MediaFind doesn't do. The agent and developer pages told people to pip install mediafind, but the package on PyPI is a placeholder: the agent guide now starts from Settings → Connect to AI agents, and the SDK pages say plainly that it isn't publicly installable yet. The free plan no longer promises "15+ search modes" (faces and logos are Pro), and Pro no longer lists duplicate culling (exact duplicates are free; near-duplicate and quality culling are Pro).
  • Create is called Create everywhere. Its window, the sidebar and the tutorial still said MediaCreate, a name the website retired, and the tutorial pointed to a Find ⇄ Create toggle in a top bar that has none. Create opens from the sidebar.
  • The ffmpeg and folder-picker messages are true on every system. Quick edit's missing-ffmpeg message says how to install ffmpeg on Windows too, and that only a source install needs one: the Mac, Windows and Linux apps include their own. In a browser, choosing a folder said "the desktop app can" read its path, though the button just pressed couldn't; it now points to File ▸ Add Folder… in the Windows and Linux apps, which can.
  • Demo mode stays on the example clips. Demo mode opens every Pro feature so you can try them on the clips MediaFind ships with, but it opened them for everything else in your library too. It now turns on only while your library holds none of your own media and nothing is being added to it, and nothing else can be added while it is on: Create's import, live meetings, restoring a backup, and importing or syncing a library are refused like the other ways of adding media. A demo mode left on over your own files ends the next time MediaFind starts, and takes its example clips with it.
  • Demo mode's editor exports stay on the example clips too. Exporting a Create project to an editor (FCPXML, EDL, OTIO, AAF), importing a timeline back from one, and exporting a selection are Pro features that read whatever files a project or a selection names. In demo mode they now take only the example clips and what MediaFind made for the project: captions, titles and generated stills.
  • Deleting your face data, or a live meeting's recording, no longer needs Pro. Faces are found by default during the free trial, but once it ended "Forget all faces" answered with an upgrade prompt, and the People page didn't even show the button: the only way to delete that face data was to subscribe. It now works on every plan, and the button is there whenever face data is stored. A live meeting still going when Pro ran out could be neither saved nor discarded, and pressing Stop left a microphone or call capture recording. Stop now always ends the recording and Discard always deletes it; only saving it as a meeting needs Pro.
  • An assistant connected over MCP can no longer delete the conversations you imported. Its memory tools use the same library names as the Conversations page, which files imports under "personal" by default, and memory_reset deleted whatever library it was given. It now clears only a library whose every conversation the assistant saved with memory_save, which marks each one it creates. A library holding anything else — your imported ChatGPT or Claude history, notes, or memory an assistant saved before this version — is refused and left as it is, for you to reset in Settings → Memory. Settings and the "MediaFind for AI agents" page now name the memory tools and say what they can reach.
  • The free plan's file cap holds for libraries you import or sync. Importing a library export, or syncing with another Mac, added its files without counting them against the cap. On the free plan, an import or a sync that would take your library past the cap is now refused before anything is written, and the message says how many files you have, how many it would add, and the cap. Files already in your library still bring their tags, notes and collections, conversations take no file slot, and Pro is unchanged.
  • mediafind restore refuses a backup made by a newer MediaFind. Restoring one put a library this version cannot open in place of yours, and then changed it with this version's schema. It now stops before touching either file and says why, as the app's own restore already did.
  • The tutorial's Meetings section no longer offers a video that isn't there. Its *Watch the demo* button played a file the website never had, so it is gone, and the tutorial's Pro paragraph points to the full list on the pricing section instead of naming four features of it.
  • Home has a heading of its own, "Your library". It shared Find's "Find the moment", and screen readers announced Home as that page.
  • Add Media starts in the local folder field, the free way in, not the Pro URL box.
  • The OCR switches' tooltips say whether this computer can read on-screen text, from what MediaFind finds when it looks for its OCR engine, instead of a fixed claim about which download includes it.
  • "Forget all faces" now deletes the face images too. It removed every face and person from the index, but each face MediaFind had cropped from your videos and photos stayed on disk as an image in the faces folder of its data directory, so the most sensitive thing it keeps outlived the button meant to wipe it. mediafind people --forget only printed an rm -rf hint. Both now delete those images as well, including any an earlier re-scan had left behind, and say how many; if some cannot be deleted, they say where they are. Nothing outside that folder is touched.
Changed
  • The sidebar is built around MediaFind's four parts: Find, Meetings, Memory and Create, each at the top level after Home and Add Media. People & Voices and Organize stay under Library, Quick edit and Cleanup under Tools, and Knowledge & Insights sits in a group named Insights (it was "Agents"). No two entries share an icon any more.
  • Memory has one name. Conversations is Memory now: in the sidebar, on its page and in Add Media ("Memory — chats & notes"). Its named sets of conversations, such as personal and work, are *spaces*, so they can't be taken for your media library; the API and the command line still call them libraries. The Home rail of moments from past years is On this day, the search badge for a word you corrected once says Learned spellings, and Ask's New conversation is New chat.
  • Quick edit and Create. The ffmpeg Video Editor is Quick edit, for trimming, converting or resizing one file, pulling out its audio, grabbing a frame or making a GIF; it points at Create for a timeline of several clips, and Create's Projects page links back to it. Create no longer shows a Generate tab: local still generation is off unless the server enables it, and there the room only said so.
  • Each reel says what it makes. Find's Search reel stitches the moments matching your search into one video; the player's Highlight reel picks one file's key moments; Organize's highlight reel of #tag does that across a tag's files; and Home's Rough-cut plan (it was "Rough Cut") is a plan, a cited moment for each beat of your story with alternates, with nothing cut or rendered.
  • The first-run wizard asks for the transcription model first, then the AI assistant: indexing needs the one, and the other is optional.
  • The Mac app says which macOS it needs, and everything in it runs there. Earlier versions did not say, so they opened on any macOS, yet the engine behind their on-device Ask, summaries and meeting notes, and their voice detection, were built for macOS 26, and search and faces need macOS 14; on an older Mac those failed when used. The app now asks for macOS 14 (Sonoma) or later, and Finder says so on an older Mac. Everything inside it is built for macOS 14: a release rebuilds the libraries it compiles from source for that version and stops before signing if any file in the app needs a newer one. The AI engine is also built for every Apple Silicon Mac rather than for the release Mac's chip: earlier builds used instructions an M1 does not have, which could stop on-device Ask on an M1 when it ran on the processor. Content Credentials (C2PA) checks are off in the Mac app for now: the checker it carried came from the release Mac's Homebrew and needed macOS 26.
  • The Mac app is about 30 MB smaller. It no longer carries a type checker, a coverage tool or the Tk toolkit, which nothing in the app uses, or duplicate copies of MediaFind's own modules.
0.1.55
Added
  • The text MediaFind read off the screen can now be downloaded. The "On-screen text" panel under a video — the burned-in subtitles, slides and signs it read from the picture — has a Download of its own, as TXT, SRT, VTT, Markdown or CSV, with the same "redact PII" option as the transcript below it. Subtitle timings are worked out from how often the file was sampled: text held across several frames becomes one cue, and no cue covers the stretch where nothing was on screen. Keyless and on-device like the transcript download, and saved as <name>.on-screen-text.<ext> so the two never get mixed up. mediafind export-ocr <file> prints the same thing from the terminal.
Fixed
  • Clicking a voice in Library Insights opens its clips again. If you had renamed a voice, its bar was still sized by everything that person says — but clicking it opened an empty page saying there were no clips for them. It was looking them up by the name you gave them rather than by the voice itself.
  • The activity calendar shows the right days outside the Americas. Everywhere east of the UK, every square in the "by day" heatmap carried the previous day's count, the first active day looked empty, and the last day's files were missing from the grid entirely — a library with activity on a single day showed a blank calendar.
  • Two downloads stop waiting when there is nothing left to wait for. If MediaFind restarted while the visual-verifier model was downloading, or while visual evidence was being prepared, the dialog sat on "Downloading locally…" for as long as the page stayed open, with the button stuck greyed out.
  • "Add all to selection" saves the moments you are looking at. After "Separate nearby moments" split them apart, it still saved the merged set — fewer moments than the page was showing, with no sign anything differed.
  • Exporting follow-ups to a calendar says when it fails. If the library was busy with an import, the button greyed for a moment and then did nothing at all — no message, no file, nothing to tell it from a cancelled save.
  • The assistant stops congratulating itself for tagging nothing. When every file in a delegated tagging task had left the library, it reported "✓ Tagged 0 file(s)".
  • The follow-ups list can no longer show decisions as action items. A refresh landing at the same moment you switched tabs could paint one list through the other's renderer — leaving rows with no text, or worse, a tick-box that completed an unrelated action item.
  • A failed "Detect speakers" can be retried. When detection failed, the message replaced the panel that held the button — "⚠ Speaker detection failed — try again" with nothing to click — until you opened another file and came back.
  • Renaming an open collection keeps it working. Renaming it from its own header left the panel under the old name, so removing a file from it failed every time.
  • An open collection drops a file you remove from the library. It kept listing and counting the file next to the lower count in the list and sidebar, and removing it from the collection then failed with "'…' is not indexed". A file that is already gone now just disappears from the list, with a note saying so.
  • A file's collection chips stay current. After you renamed or deleted a collection, added the open file to one, or removed it from one, the Player's chips kept the old list, and a chip for a renamed or deleted collection did nothing when clicked. The chips now update, and a stale chip says the collection was renamed or deleted.
  • Creating a collection that already exists says so, however it is spelled — "Road Trip" with two spaces, or "STRASSE" beside "Straße" — in the app and in mediafind collections create, instead of "Collection created."
  • Adding a file to a collection from Home stays on Home. "+ Collection" in a file's drawer jumped to the Player and swapped the file the Player's tools act on — even when cancelled — so a tag, a re-transcribe or "Remove from library" there could hit a different file from the one playing.
  • "Auto-categorize all" no longer hides the suggestion for good. Clicking it hid "N files have no category" permanently, even when nothing could be categorized, so files added later were never flagged. It now says how many it categorized.
  • Removing selected files warns about hidden ones after a search. Once you had searched, a pick hidden by a Home filter no longer triggered the "hidden in the current view and will also be removed" warning.
  • Back and the breadcrumbs in a Home folder stay inside your library. Back from a folder you indexed climbed into its parent directories, and the path above the grid left out the folder itself.
  • Saving a search under a name you already use asks first, instead of silently replacing the saved search.
  • The newest search wins on Find, including "count …" and "how many …" questions. Asking one of those while another search was still running (or the other way round) let whichever answer came back last replace the other, under the words you had just typed. Escape and New search now also stop a search that is still running, so its results no longer appear under an empty search box.
  • "Count …" answers count inside the folder you picked, and say which of your other filters they didn't apply (dates, speaker, person, channels) instead of counting across the whole library as if they had.
  • New search starts from a clean slate. It cleared speaker, person and channels, but kept the scope, sort, folder and dates, so the next search, including one typed on Home, could say "No matches" under filters the button said it had cleared.
  • The Folder filter and the scope pill no longer contradict each other. With a scope set, picking a Folder lit it up as applied while the results stayed in the scope; now whichever you pick last replaces the other.
  • Picking a speaker or person re-runs the search, as sort, folder and dates already did. Before, nothing changed until "Show more results", which then started over and showed fewer results.
  • A saved-search suggestion runs the saved search, with its folder, dates and channels, instead of just its words across the whole library.
  • Opening a file from a tag's list plays that file, not the one that was already in the player.
  • Action and name pages list every file their chip counts. A label was read as a typed search, so "Playing video games (2)" found nothing, "Taking a photo (2)" found only the photo, and the name "Zoom" turned into a camera-zoom filter. Pages also stopped at 12 results without saying so, so "Acme Corp (16)" showed 12.
  • A brand's page shows that brand alone. The "Nike" page also listed "Nike Jordan" clips, and "X" listed Xbox. "Export all as reel" would have cut those clips too.
  • A slow brand, action, name or voice page no longer lands on the one you opened after it.
  • An open tag's file list drops a file you remove from the library. It kept listing the file, and opening it showed an empty player. Opening a removed file now says it is no longer in your library.
  • The Brands & logos and Actions lists show what was found even without the visual add-on. They went blank when the visual add-on wasn't installed, although Knowledge & Insights listed the same brands and actions. A note now says that only detecting new ones needs it.
  • Category chips on Home stay in step. A slow response could leave one chip pressed over another category's files, and after recomputing categories Home kept chips for categories that no longer existed.
  • A large library no longer freezes the app every few seconds. With thousands of files, each status refresh — every 5 seconds during an import — stalled every page for about 9 seconds while it redrew the Knowledge & Insights map out of sight. The map now loads only on its own page, and on a big library it shows the 500 most connected files and says so ("500 of 5,519 files").
  • A busy library no longer empties a category on Home. A category opened during an import could show no files until the next refresh, and a busy category list hid the chips while a filter stayed on. Home now says the files didn't load, with Retry and Show all files, and the chips stay.
  • Keyboard focus stays where it was when Home refreshes. A refresh sent focus from a file, folder or checkbox back to the top of the page, and so did closing a file's details after a refresh.
  • Arrow keys on a folder in icon view no longer open folders in the list view.
  • A file removed elsewhere is unselected on Home, instead of staying selected as "hidden in this view" and making Remove fail. "Add to collection" now also says when some selected files are hidden in the current view.
  • Smaller fixes in Knowledge & Insights and Organize: focus stays on a facet heading you expand; creating a collection or a tag that already exists (in any capitalisation) says so; hidden map labels stay hidden when the map reloads; the People rail counts faces as faces, not files.
  • "Retry failed files" retries the files that failed. If you had picked other files in the meantime, it quietly indexed those instead, while showing the failed folder.
  • Retry re-indexes exactly the files that failed. After a run over several picked folders it re-indexed only the first folder and then hid the failure. After you removed a folder, its Retry brought the whole folder back. Removing a file or folder now also drops it from the failure report.
  • Re-adding a folder that's already indexed says so. It said "No supported media found in that folder." whenever nothing new was indexed, and also hid files that failed again. It now says "Nothing new to index — N files already up to date", or reports the failures.
  • Opening Settings no longer re-ticks "faces" on Add Media. Unticking face detection for a run was undone by any visit to Settings, so the next index detected faces anyway.
  • Clicking Index twice no longer indexes the folder twice at the same time. The second request now waits for the first to finish and only picks up files added since.
  • "Remove media" lists the folders you actually indexed. On a large library the list was cut to 200 entries from the wrong end — keeping deep subfolders and dropping the top-level folders you added — with no sign it had been cut.
  • Saving a preference says why it didn't save. It always read "Could not save preferences.", even when the reason was specific, like a setting locked by an environment variable.
  • Multi-device sync goes to the folder you typed. A sync folder starting with ~, or a relative one, was created inside MediaFind's own working folder, while "Sync now" still said "Synced", so other machines never saw the export. ~ now expands to your home folder, and a relative path is refused with a message saying why.
  • Preferences show up as soon as Settings opens. They stayed blank until a full database check finished, and saving in that window failed with "'embed_lang' must be one of …". Save now waits for the preferences to load.
  • Demo mode no longer claims you have a license. Settings said "✓ Pro — unlimited indexing", hid the key box, and offered to remove a license nobody had entered. It now says demo mode is on, and the key box stays available.
  • Relinking moved files handles folder paths exactly. A new folder typed with a trailing slash stored paths with a double slash, the old folder …/Trip also matched …/Trip2, and a relative new folder was accepted. Each is now handled.
  • "Add to Claude Desktop" explains a symlinked config file instead of saying only "Could not add to Claude Desktop."
  • Turning off "Access from your phone" locks phones out at once. Until you quit MediaFind, the token kept working while Settings showed the switch as off. A token changed in Settings was shown but not enforced. New token replaces the token and immediately locks out the old one. The card now says when the server is still open to your network, or when it needs a restart to open.
  • The phone-access card says what the token allows: using MediaFind fully, over an unencrypted local-network connection. Where the switch can't apply (MediaFind started with mediafind serve, as on Windows and Linux), the card says so instead of showing an address that doesn't answer.
  • /metrics needs the token from other devices, like every other page except the pairing page and /health. It showed anyone on the network how many searches ran and how long they took.
  • The activity dock no longer blocks taps along the bottom of the screen, and Create's "Delete project" button is visible on touch screens and to the keyboard.
  • Next and Previous in the player stay on your search. Opening a person's, voice's, brand's or action's page from a file swapped out the list the player was stepping through: pressing Next then jumped into that person's clips ("2 of 20", still labelled with your search), and clicking a result back on Find could open one of their clips instead of the result you clicked.
  • A saved search's snapshot time is in your time zone. It showed the stored UTC time as-is, so a search saved in the evening in California read as the next morning.
  • Confirming an out-of-date voice match says so. After merging voices or people, a suggestion naming one that no longer exists greyed its buttons and did nothing; it now explains and refreshes the list.
  • A voice match links the person it names, even after "Re-group faces". Re-grouping renumbers everyone, but the suggestion cards, Identities, the open file's face chips and Find's person filter kept the old numbers. Confirming "SPEAKER_01 → O'Brien" linked the voice to whoever was now Person 1. They all reload now, and a person filter that no longer means anyone is cleared, with a note saying so.
  • A voice linked to someone without a name shows the link. It looked unlinked, and 🔗 offered to link it again with the first person already picked, so one click moved it. It now reads "→ Person 3" and offers Unlink. The Link dialog lists everyone, can be searched, and tells two people with the same name apart.
  • A slow page no longer lands on the person or voice you opened after it. A voice's clips could fill another voice's page, whose rename then renamed the first voice. A person's page could be replaced by an error for someone else, or closed.
  • Libraries with more than 500 people keep the Notable filter. Its tab and the duplicate-celebrity bar disappeared when no recognized person was among the first 500, and came back only after "Show more".
  • Switching to Notable unticks the people it hides, so Merge can no longer combine people you can't see.
  • Moving faces that changed since the page loaded says nothing moved. Split and "Move to existing…" reported success when a re-index had already replaced those faces.
  • Smaller People & Voices fixes: renaming a voice from its own page shows the new name, renaming a person updates Identities, and detaching a voice removes its "→ person" link. Suggested voice matches no longer show under a person's page, the "Show more · 1 more person" tile that never loaded anyone is gone, face crops are buttons only while splitting, and Back returns focus to the tile you opened.
  • "Ask a follow-up" shows the overview's answer once. Following up on an AI overview put its question and answer into the Ask thread twice.
  • Starting a new conversation while an answer is still coming leaves that answer behind. It used to land in the new thread with no question above it, and your next question was treated as a follow-up to the conversation you had just closed.
  • Ask takes one question at a time. Pressing Enter while an answer was still being written sent a second question beside the first: the answers landed in the wrong order, the two questions went into two separate conversations, and after a reload only the first came back. The second question now waits in the box.
  • A search on Find no longer strands the answer Ask was writing. It appeared without its question, or not at all. It now shows under its question when you go back to Ask.
  • "Ask a follow-up" can be clicked twice. The second click emptied the thread and lost the conversation, so the next question started a new one.
  • An AI overview no longer appears over a newer search, or on a Find page you had just cleared. It is also skipped when a speaker, person or date filter narrows the results, since it can't follow those filters, and it now answers within the folder you picked.
  • Ask keeps a conversation it couldn't load. One failed load (a busy library) forgot the conversation for good. It now says so and offers Retry.
  • Opening a file no longer re-scopes the conversation you're in. A file opened from Recent files quietly narrowed your next follow-up to that one file.
  • A follow-up on an AI overview asks within the overview's scope, not within folders left picked on the Ask page from an earlier conversation.
  • A cited file removed from your library says so, instead of blaming the browser's codec support.
  • Ask follows Settings ▸ Ask backend again. The hidden backend picker had no entry for the default setting, so every question skipped your choice until a reload.
  • Each clip page keeps its own clips and its own Export selection. After opening a voice's or a brand's clips and then a person's, ticking a clip on the person's page lit up the hidden page's Export bar instead of the one on screen. Going Back to the voice, its clip played one of the person's clips instead, and its Export offered the person's selection. The player also labelled a person's clips with the previous page's name.
  • "Find in search" filters to the person you picked, however many people you have. The person filter lists the 300 people seen most often; for anyone else it said "Search is filtered to …" while the search was not filtered at all.
  • Expanding a folder with the arrow keys keeps your place. In the library's list view focus was lost, so the next arrow key scrolled the page instead of moving through the folders.
  • Stepping onto a photo in the player stops the video. Pressing ▶ Next from a video result onto an image result left the video playing behind the photo, with its player and tools still showing — and a note added then was saved on the video.
  • A highlight reel stays with its file. Its moments stayed on screen after you opened another file, and clicking one loaded the first file back into the player while the transcript and notes below stayed on the second.
  • A busy library no longer shows the previous file's details. Opening a file while MediaFind was updating the index showed "Indexing in progress" above the last file's location, size, tags and chapters.
  • "Reattach notes from another version" says when nothing was reattached, and why. When the two files couldn't be lined up — one without visual frames indexed, say — it just reported "✓ 0 reattached".
  • A file that has moved says so. Playing a file that is no longer where it was indexed — moved, renamed, deleted, or on a drive that isn't connected — blamed your browser's video support and offered a download that failed too. It now says the file has moved and offers to relink it.
  • A result that couldn't open no longer scrolls the next file's transcript. The next file you opened jumped to, and highlighted, a line at the same time as that result.
  • Speakers that couldn't be loaded aren't reported as silence. The panel read "No speech detected in this recording." above a transcript full of speech; it now offers a retry.
  • A moment from a highlight reel, Knowledge & Insights or a workflow opens its own file. Clicking one played it while the page kept showing the file you had open, so a note added then was saved on the other file. Notes now go to the file the page shows; if the player holds a different one, the note says so and keeps what you typed.
  • A slow highlight reel no longer lands on the next file you open, and a clip's share or export status stays with its own file. Opening another file also clears the last one's tool results, such as its speakers' talk time or its protection bundle's path.
  • An open transcript edit survives other changes. Cancelling or saving another line, or adding a tag, threw away what you had typed; the line now reopens with your text and cursor where they were.
  • Clicking a transcript line while its file is still loading plays from that line, not from the start.
  • ↑ and ↓ move through transcript lines, chapters and notes again.
  • Renaming a speaker from the transcript updates the Speakers panel.
  • Turning your recordings into meeting briefs no longer forgets who was speaking. If a recording arrived already knowing its speakers — a transcript you imported, a live meeting you saved, the bundled demo clips — "Process recordings" ran the voice detector over it anyway and wrote its own labels on top. Every name was replaced by "SPEAKER_00", everywhere at once: the transcript, People & Voices, the speaker filter in Find, and the person each action item was assigned to. A recording that already says who spoke is now left alone, and if the detector does run over one, the names its lines already carried survive it.
  • A half-finished batch of meeting briefs says which recordings it could not do. "3 processed · 3 failed" was the entire report — no names, no reason, and nothing in the log either, though the reason had been worked out and thrown away. It now names them and gives the first reason ("no meeting artifacts produced for holiday.mp4 (empty transcript?)"), with every failure and its reason in the log.
  • A live meeting with nothing transcribed says so. It reported "Live meeting saved" and then "Could not load this meeting — try again", because no meeting had been saved; and when the transcriber failed on every line it still said "no speech detected". It now says nothing was transcribed and why, and where the recording was kept.
  • Meetings are dated by when they were recorded. Every meeting was dated, sorted and scheduled from the moment it was processed. A backlog processed today showed every meeting as today's, newest recordings last, and "by Friday" from a meeting last November landed on this Friday in Follow-ups and in the calendar export. The list, the brief, Follow-ups, the calendar files and the Word/PDF notes now go by the recording's date.
  • A live meeting that is still recording shows on every page. Leaving Meetings hid it completely while it kept recording. A REC pill in the top bar now shows the running time and takes you back to it.
  • Checking off a follow-up can be undone, and shows as done everywhere. It vanished with no way back, while the meeting's brief and stats still counted it as open and its calendar export reopened it. The toast now offers Undo, and the brief, stats, Copy and calendar show the item done.
  • "Open recording" opens the file it plays. The player switched to the meeting's recording but kept showing, and acting on, the file that was open before.
  • A slow brief no longer covers the meeting you clicked next, along with its export and Copy buttons.
  • Renaming a speaker updates an open meeting brief. The brief and Copy kept the old label while Follow-ups showed the new name.
  • The job dock and Operations Center show what a job actually did. An index or download run that lost files showed raw JSON as its message, in success green, and the Operations Center's progress bars never filled.
  • "Time left" comes from real progress. After a page reload a job halfway through read "~1s left" with half a minute to go, and time spent waiting in the queue counted as work.
  • An untitled live meeting exports under the time it started, not a hex id.
  • The dock's pill names the job that's running, not one that just finished.
  • One toast per finished job. Several actions showed their own "done" message and a generic "✓ complete" as well.
  • Knowledge & Insights lists the same voices in both places. The Voices row along the top showed only voices the detector had measured a voiceprint for, so a library whose speakers came from its transcripts showed "SPEAKER_00, SPEAKER_01" directly above a Speakers panel naming Cecilia Rouse, Susan Rice, Thom and Celia. Both now read the same list.
  • "Show more results" stops undoing what you just did. Picking a result type (People, Files, Transcript) and then loading more results snapped the list back to All and brought back every row you had filtered away; the "How this was found" panel came back open and stuck on "loading the retrieval steps…" until you ran a new search. Both now survive loading more, and both still reset when you search again.
  • "Export top moment as clip" cuts the top result you're looking at. It ran the query again without your scope, filters or channels, so it could cut a file the list didn't even show. It now cuts the first result on screen.
  • A month named without a year means the last one. "revenue in november" asked in September searched a November still to come and found nothing; it now searches last November, as "since november" already did. "after November 1" works the same way.
  • Find says what it read into your words, and lets you search them as typed. "screen recording" quietly searched "screen" among audio files, and "landscape" became a photo-orientation filter with no words searched at all. The results now show "Interpreted as …" with a "Search the words as typed" link.
  • Searches like "audio files" list files, not weak guesses. Every file matched, yet the page said "No strong matches … closest guesses" and drew each file as an empty transcript line.
  • The Transcript tab keeps lines that name their speaker. In a library where speakers were detected, every transcript line was filed under People, and the Transcript tab disappeared.
  • Result-type tabs filter the results on screen. After you opened a voice's or a brand's clips and came back to Find, the tabs hid and showed rows by the other page's list, and a shadow card's Dismiss could act on another item.
  • The results header names the channels you searched. It always said "visual + transcript". A Sounds-like search also no longer says "No strong matches", and its results show their time.
  • Stepping through hits in the same recording no longer reloads the page around them. Prev/Next between moments in one file re-fetched the file, its transcript, its notes, its speakers and its frames every time — the summary blanked and the details line flashed "Loading…" on each step, and while indexing was running the reload could time out and replace content that was already correct. It now just moves to the moment.
  • Cleanup's Compare button says something when a group has gone. Opening a duplicate group that had since been ignored, quarantined or rebuilt away did nothing at all — no dialog, no message — because "no such group" was reported as success.
  • Cleanup's reclaimable figure covers your whole library. It was summed over the first 10,000 duplicate groups, so a larger library was shown a group count and a reclaimable size that described different sets of groups. It is now one measurement over all of them, and a much faster one.
  • Alternate takes say "try again" while indexing, instead of "none". A library the indexer was writing to reported that it had no alternate takes at all. Every other panel already offered the retry.
  • Opening a Create project shows the first frame again, instead of a black stage. Opening one runs seven redraws at once, and six of them were handed a video that had not finished loading yet — so the redraw that got there last waited forever for a frame that was never coming, and the preview stayed black until you scrubbed. During playback the same wait froze the picture while the playhead and audio carried on.
  • The Video Editor's trim boxes accept a decimal. Typing 12.5 into the In point put 12 in the box on a good day: the box was rewritten on every keystroke, which swallowed the decimal point and pushed the digits after it to the front. On a 72-second clip 12.5 came out as 72 — an empty selection at the very end of the video — and Export clip would have cut that.
  • A Video Editor export that fails says so. Export clip reported "Trimmed clip done." when no clip was written — for a range past the end of the file, or when ffmpeg failed.
  • A moved source file isn't reported as a missing ffmpeg. Opening the Video Editor on a file that had moved said ffmpeg wasn't installed and disabled every export; it now says the file moved and offers to relink it.
  • Leaving the Video Editor stops its preview. A previewed selection kept playing, on a loop, on whatever page you went to next.
  • Video Editor times are exact, and show hours. A selection could read a tenth of a second short (In 2.1 to Out 5.3 read 0:03.1), and past an hour times read "60:00.0" beside the player's "1:00:00".
  • Typing an Out point is no longer rewritten mid-number. With the In point at 10, typing 20 into Out turned into 10.050 — a 0.05-second export.
  • Share clip says why its bundle has no clip, and reveals the bundle. It blamed a missing ffmpeg every time, though a redacted bundle deliberately leaves the original clip out, and its "open share.html" link couldn't open from the app. It now offers Reveal in Finder, and a cancelled share is no longer reported as a failure.
  • Chapter lists download as files an app can open. They saved as .youtube and .ffchapters, now .youtube-chapters.txt and .ffmeta, and a file without chapters says so instead of saving an empty file.
  • Downloading more than 25 found videos explains the limit instead of failing with a validation error.
  • Notes can be reached from the keyboard. Only their delete button could take focus.
  • Per-clip buttons dim when a Bin clip is added. Adding one clears the timeline selection and says so, but Split, Duplicate, Nudge and Remove clip stayed lit; clicking them did nothing at all.
  • Delete removes the clip you picked, not the one that slid into its place. The timeline held a selection by position. After a cut in the Transcript room or ⌘Z, the highlight and Delete moved to whichever clip now sat there, while the label still named the one you picked. A selection whose clip has moved is now cleared.
  • ⌘S no longer splits the selected clip. Only a bare S does.
  • The Transcript room reports a cut at the start or end of a clip. Cutting a clip's first or last line said "No transcript ranges matched", although the clip had just been shortened.
  • Opening another project starts with a clean timeline. The last project's "Selected: …" text, lit clip buttons, zoom and scroll position came along with it.
  • Smaller Create fixes: going Back right after renaming a project lists the new title, and audio clips and music beds are named after their file instead of "m1".
  • A duplicate scan reports what it found in real units. "Found 0 duplicate group(s) — 0.00 GB reclaimable" was always in GB, so a scan that turned up 4 MB to reclaim said 0.00 GB directly under a Reclaimable tile reading 4 MB.
  • Quarantined files stay out of later cleanup scans. Scanning again brought photos and videos you had already quarantined back as low-quality or near-duplicate suggestions — and a near-duplicate group could then offer the quarantined original as the copy to keep, with the only copy still on disk ticked for quarantine. Quarantining a selection that would leave no copy on disk is now refused.
  • Cleanup tells you when it didn't do something. "Keep best, quarantine rest" did nothing at all when the copy it would keep was missing or the group had changed since the list was drawn, and quarantining or ignoring from Compare — opened from a search result or Home — reported only on the Cleanup page. Skipped files now say why.
  • A scan's summary counts what the tiles count. A Pro scan's "Found 5 duplicate groups" also counted blurry and low-resolution photos, under a Duplicate groups tile reading 1.
  • Batch Apply uses the rule you can see. After previewing one keeper rule and switching to another, Apply still used the first; a failed apply showed nothing.
  • Workflow results survive playing one of their moments. Clicking a moment in a Research Quote Pack or Rough Cut closed the dialog, and reopening the card showed an empty form; closing it mid-run lost the result too. The card now shows its last run. Run also waits for edited inputs to be checked, instead of failing with "request failed (HTTP 409)".
  • Cleanup goes back to duplicates when Pro ends. Left on the Low quality tab, it showed duplicate groups as nameless low-quality files.
  • Importing a chat no longer inflates Home's transcript count. Home counted a conversation's turns among the "timestamped transcript segments" of your video library while leaving the conversation out of the file count in the same sentence, so importing a six-message chat turned 44 segments into 50.
  • A quoted message opens where it really is after a chat is re-imported. Messages in text chats and notes are numbered by position, so after re-importing an edited copy, a citation or search result opened and highlighted a different message. The page now checks that the message still says what was quoted. If it moved, the page shows where it is now. If it's gone, the page says so.
  • The activity dock counts files the way the header does. A library of six imported chats read "6 files · 0 segments indexed" next to "nothing indexed yet".
  • Saved searches over conversations show the conversation's title, not an id like "personal~a9b8c7d6-budget".
  • Conversations' description matches separate-library mode. It said conversations show up in Find results, which isn't true when each library keeps its own database.
  • Times typed at the start of a chat line are read on your computer's clock. A pasted "[2025-09-01 18:30] Dana: …" showed as 11:30 AM in California. A time with a zone or offset still names that zone.
  • Opening a conversation from Find with the keyboard keeps your place. Focus moves to the highlighted message or to the conversation's title, instead of falling back to the top of the page.
  • A busy library no longer reads as an empty one. While an import held the write lock, the status the page polls came back "try again in a moment" — and the page drew that as a library with nothing in it: "nothing indexed yet" in the top bar, the sidebar counts at zero, the first-run "Add media" card popping open over a full library, and the speaker filter quietly resetting to "any" so the next search returned unfiltered results. It now keeps what is on screen and picks the real numbers up on the next refresh.
  • Dragging a clip's left edge trims its head. It did the opposite of what the timeline previewed: dragging the left handle inwards — which draws the clip shrinking — asked to *extend* the head instead, so it failed and snapped back on every clip added from the Bin; dragging outwards, previewed as longer, shortened it.
  • A refused timeline edit says why. Every one of them reported a bare code — "Move failed: move 422" — while the server had sent the actual reason ("frame 75 falls inside an existing clip"). The reason is shown now, and no longer stretches the toolbar down the page when it is a long one.
  • Cutting a line in the Transcript room takes it out of the sound too. Adding a clip places its picture and a copy of its sound, but a transcript cut (clicking a line, Remove silences, Remove fillers, Keep speaker) cut only the picture. The sentence still played, and everything after it was out of sync. A cut across a split point also left its second half in.
  • Footage at 25, 24 or 23.976 fps can be split, trimmed and cut in a 30 fps project. Split and trim refused many frames ("cut cannot be represented exactly in source timebase"), and Remove silences failed with "request failed (500)". Split and trim now land on the nearest frame the footage has, a few frames at most, and transcript cuts remove exactly the frames they should.
  • A rendered video keeps the gaps you left on the timeline. The picture joined its clips back to back, so after a gap it ran ahead of its sound, and a clip placed after the start began at 0.
  • Dragging a clip later along its track drops it where you let go. It landed one clip-length further on, after a stretch of black or after the next clip.
  • The caption track is called "Captions". It showed the internal name __captions__ on the timeline, where every other clip reads normally — and it is no longer treated as footage anywhere else either: "Generate editing proxies" tried to transcode it and reported a failure against it on every captioned project, and the Transcript room counted it as a second source, so a project with one video asked "which clip do you want to edit?" and offered a caption track that can never have a transcript.
  • Script to screen says so when it adds nothing. Matching your lines and then clicking "Add matched lines" on a brand-new project placed none of them — the footage it matched isn't in the project yet, so its frame rate is unknown and each line is skipped rather than guessed at. It reported that as success: "Added 0 clips — see the Timeline room", in green, pointing at an empty timeline. It now says what happened and what to do about it. Also: "1 line needs footage".
  • "Assemble as a new project" opens the project it made. It drew the new cut in the project you had open, while the title, Render and every later edit still went to the old one — so trimming or removing a clip could change a clip you couldn't see. When none of the footage could be placed it now says so instead of making an empty copy.
  • The Storyboard says so when it adds nothing. Assembling beats into a project that couldn't place their footage reported "Appended 0 clips" as success.
  • "Duck under speech" ducks the music, not the dialogue. Ticked, it lowered your dialogue whenever the music bed played, and with no music bed it reported ducking on while doing nothing. It now applies to music beds and waits until there is one.
  • Edits stay in the project you are working on. A slow reply from a project you had just left (a music bed, an undo, a rename) could replace the timeline of the project you opened next, and your next Delete then removed a clip from that project.
  • Building editing proxies no longer undoes your edits. A clip you deleted while proxies were being built came back when they finished, and building two proxies at once lost one of them.
  • The shortcuts list no longer lets keys through to the timeline. With it open, Backspace deleted the selected clip hidden behind it.
  • Saving a version under a name you already used asks first, instead of silently replacing that version. Restoring also says when it will replace the previous pre-restore backup.
  • A running render or proxy can be cancelled. The job card's × came back on the next update and nothing in Create could stop the job. The × now cancels it after asking, and the card no longer shows its percentage twice.
  • An audio file added from the Bin goes on the audio track. On the video track the render asked for a picture it doesn't have, and the whole render failed.
  • Captions say when they are out of date. After a clip is removed or moved, captions still showed its lines at the old times. The Captions and Deliver rooms now say so before you render.
  • Gain applies on top of "Normalize loudness". Normalising undid the gain in the render, while the preview still played it.
  • Adding a title no longer moves the others. A title added earlier on the timeline pushed every later title, lower-third and brand logo along by its length, and one starting during another title was refused.
  • Versions shows edits that don't move a clip. Comparing with a saved version after a grade, a reframe, an audio change, a transition or a caption restyle said "no changes".
  • Undo survives tidying up Versions. Versions listed Undo's own internal steps, each with a Delete button, and deleting one stopped Undo from reaching any earlier step. Restoring "pre-restore" also no longer promises a backup it doesn't make.
  • The Undo button lights up after an edit made in any room. It stayed greyed out after Color, Reframe, Text, Captions, Audio and other rooms' edits, though ⌘Z worked.
  • A refused music bed or assembly says why. Those messages read "[object Object]".
  • A generated shot is still there when you come back. Leaving the Generate room while a still was generating lost it, and "Generate shot" could be clicked again.
  • Auto-Reel offers only sources it can mine. Its Source list included the caption track, title and logo images, music beds and generated stills, each failing with "transcribe it first".
  • The caption style preview no longer draws a second caption over the real one.
  • Project cards show a readable date. The Create home printed the raw stored timestamp — 2026-09-16T08:05:14+00:00 — where the rest of MediaFind says "Sep 16, 1:05 AM".
  • Small counts read properly. "1 categories" on Home, "1 files" in the search scope picker and in a category's label, and a meeting with one action item no longer claims a tick against "0 decisions".
  • "Redact PII" stays ticked. In Settings → Memory, deleting a conversation, resetting any library or reopening Settings cleared the tick without a word, so the next Export was not redacted.
  • Model downloads no longer show a button that does nothing. A "Continue without it" button appeared under every download in first-run setup and did nothing when clicked; it now appears only if the download fails.
  • Clicking another model during a download no longer starts a second one. It re-enabled the button, and pressing it queued a second download — and in first-run setup, whichever finished second moved you back a step.
  • The guided tour leaves Enter to the button you are on. Enter on "Start tour" skipped the first step, Enter on "Skip tour" or "Back" moved forward instead, and the arrow keys jumped into the tour while its demo library was still being set up.
  • Keyboard focus stays in place in Conversations and first-run setup. Picking a library or a conversation, or reaching a model step in setup, dropped focus to the top of the page.
  • A conversation named "." or ".." says it can't be opened instead of loading forever.
  • Going Back to the tour's brands step shows the brands again, instead of an empty spotlight on Home.
Changed
  • Content Credentials are checked once per file on a page of results, not once per result, and a person's or a brand's clips work out their durations in one query rather than two per file.
0.1.54
Fixed
  • MediaFind keeps its own address when you quit and reopen it. If anything was still connected when MediaFind closed — a browser tab, a connected assistant, your phone on the same Wi-Fi — the address it serves on stayed occupied for a while afterwards, and the next launch quietly moved the app somewhere else. The window still worked, so nothing looked wrong, while an assistant set up to reach MediaFind, a phone, or a bookmark pointing at the usual address stopped finding it. MediaFind now takes its usual address back. A second copy of MediaFind still gets an address of its own, as before.
  • The Memories panel on Home no longer keeps the rest of the page waiting. Building its "on this day" and "*someone* over the years" cards asked the database for one person's files at a time, so a library with 28,031 people cost 28,031 queries — about three minutes of database work for a panel that Home re-builds on every visit, and every other request the page was waiting on queued behind it. It now reads every person's files in one query: on that library the panel builds in 0.17 s.
  • When the on-device model says it can't answer, Ask says so too. Asking your recordings something they don't discuss, with the local model turned on, produced a confident-looking answer anyway: the model replied that it could not find an answer, that reply was thrown away, and the quoted fallback — which has no way to say "I don't know" — stitched one together from whatever was nearest. Ask now passes the model's answer on: "The closest moments don't answer that.", with the moments it read listed underneath so you can judge for yourself. It claims nothing about the rest of your library, only about what it read. A model that crashed, timed out or came back empty is not treated as a refusal — that still falls back to the quoted answer, as before — and a reply that hedges before answering ("I don't know. In 2019.") is read as the answer it is, rather than thrown away with it.
  • Ask says it doesn't know, instead of quoting the nearest thing it found. When an answer rests on your conversations, Ask holds it to the standard the Conversations page uses — answer only from messages that cover the question — but that standard was switched off the moment a single recording appeared among the sources, which on any real library is always. Asking your library something your chats do not cover ("what is sam's phone number?") therefore came back with a confident-sounding line stitched from whatever was nearest. It now says "Your conversations don't cover that." and shows the closest messages, as the Conversations page does. A recording that *does* answer keeps its answer: the test is the strongest source of either kind, and it runs only after the on-device model has had its turn — so a question the model can answer from a message the search ranked poorly ("how many cats does sam have?") is still answered.
  • mediafind ask, the MCP tools and the SDK answer from your conversations too. The Ask page learned to search your imported chats beside your recordings; every other way of asking — the command line, the tools an assistant connects to, the Python SDK and the older /api/ask — went on answering a question about your conversations from the videos that crowd them out, because they all ask through one shared function that had been left behind. That function now runs the same pass, so the answer is the same wherever you ask it from, and asking about one file still stays in that file.
  • No more naming a cartoon after a celebrity. The face detector hands back a descriptor for anything it fires on, including the things it is *unsure* are faces at all, and those descriptors were matched against the celebrity gallery like any other. On a real library that produced 19 separate people named "Harrison Schmitt" and 4 named "John Huston" whose pictures are Homer Simpson, Thanos, a Minion, a sleeping baby, an elephant, a hat and a steering wheel. Naming a face now requires the detector to have been reasonably sure it *was* a face; on that library it dropped 20 of the 23 bogus records, including every "John Huston", and changed nothing about the other 130 people it had named correctly. This is not a picture-quality rule and does not hold back old, low-resolution or webcam footage — a face filmed at 144p scores just as confidently as the same face at 720p. A name already given out this way is taken back on the next Faces & People → Identify notable faces; a name you typed yourself is never touched, and a face that misses the bar still clusters, still shows up and can still be named by hand.
  • One Elon Musk, not four. Recognizing a famous face never merged anything, so every time clustering split someone across a change of pose or lighting, each fragment was named separately and the People page listed the same public figure over and over — a real library held four "Elon Musk" records, seven "Donald J. Trump" and twelve "Lionel Messi". Recognition now folds them together in the same pass: records matched to the same person, whose faces also agree with each other, become one. On that library it merged 135 of the 136 duplicate records and left every face, every name you had typed and every dismissal untouched. Existing libraries are repaired by Faces & People → Identify notable faces. A duplicate the old "Review" suggestion offered was already rare — it required a confident match on both records, and a fragment is by nature the odd-pose cluster matched *least* confidently, so none of those seven Trump records had ever been offered. Someone you named yourself, or an identity you dismissed, is still never merged by this pass.
  • Ask answers from your conversations, and keeps the thread. In a library that holds both recordings and imported chats or notes, every memory question was answered from the videos: the messages are a handful of short lines competing with every spoken second of the library in one ranked list, so the message that answers the question never reached the part of the search that would have picked it, and Ask quoted an unrelated broadcast instead of saying it did not know. Conversations are now searched as themselves alongside the recordings, so "where did sam live?", "where did sam move to?" and "how long did sam stay in berlin?" answer from the chats that say so. Follow-ups are read as follow-ups: "how long did he stay there?" was treated as a brand-new question about nobody in particular and searched for with no subject at all, because only a question starting with why, when or where was checked for a pronoun. A question that leans on one now carries the thread with it whatever word it starts with, and answers from the same messages the question before it did. The one widened retry no longer drags the earlier turns' file names into the search, which had turned a question Ask answers on its own into three unrelated video clips by the third turn. The Ask tab can also be pointed at Conversations the way Search can, a promoted AI-overview thread keeps the scope it was answered under, and the thread itself survives running a search or pressing Escape rather than vanishing while the conversation lived on behind it.
  • The tag panel no longer stalls the page while your library is importing. Every file that finished importing threw away the counted tag panel's work, so the page — which refreshes itself every few seconds — rebuilt that panel from scratch on every single refresh. On a library of 1,000 videos that was 1.5 seconds of waiting per refresh, and the rest of the page waited with it. During an import the panel now shows the counts it last worked out and refreshes them at a steady pace instead, so the page stays quick while the numbers keep climbing as files land. A library small enough to count quickly is untouched: its panel still updates the moment a file finishes.
  • A face or brand on screen for a whole long video is one clip again, not hundreds. A video past about twelve minutes is sampled by spreading a fixed frame budget over its whole length, so its frames are seconds apart rather than two — further apart than the gap that decides whether two sightings belong to the same appearance. Nothing ever merged: a host present through a 52-minute interview came back as 236 separate two-second clips on their face page, and a brand on screen throughout an hour-long file broke up the same way. The gap now follows the spacing each file was really sampled at, so continuous screen time reads as one seekable, exportable clip while someone — or some brand — that genuinely comes and goes still gets a clip per appearance. Short files are unaffected, and face-to-voice link suggestions still weigh each sighting on its own.
  • Asking your library with the on-device model is three times faster, and costs half the memory. The model was always run on the processor, never the graphics chip, and on an Apple-silicon Mac that is the expensive way round: the processor needs its own second copy of the model to work from. One question took MediaFind from 1.4 GB to 6.8 GB and took 3.3 seconds; the same question now takes it to 4.4 GB and 0.8 seconds. MediaFind checks first that your Mac has the memory to spare for the model you chose, and stays on the processor when it doesn't, so a small Mac is never pushed into a corner — and the Encoder device setting still pins everything to the processor if you'd rather.
  • The model is let go when you stop asking, instead of sitting in memory all day. Whatever the model weighed, it stayed there until you quit MediaFind — several gigabytes held because of one question you asked in the morning. It is now released after ten minutes with no questions and loaded again the next time you ask, which takes under a second. Once you have used the on-device model, Ask goes on using it for the rest of the session — handing the memory back never changes the kind of answer your questions come back with.
  • Opening a face from a video's page no longer bounces you back to People. In a library with more than about 500 people, clicking a face under the video you were watching opened that person's page for a moment and then dropped you back on the People grid — every time, for anyone who wasn't among the most-photographed. The same mistake quietly undid three other things: people you had ticked for a merge came unticked, the "only this person" search filter reset itself to "any", and renaming or splitting one of these people closed the page you were working on. Combining people now covers everyone you ticked, rather than quietly leaving out the ones the grid had stopped showing.
  • Emotion and sounds-like searches read the words you said. An emotion search counted any word that merely began with an emotional one, so "hello" was an angry moment, "made" and "Madison" were anger, a fund or a funeral was joy and "crypto" was sadness. Its queries were read the same way: "the garage" searched for anger, while "people arguing" and "kids giggling" found nothing. It now matches a word and its own forms, so those lines drop out and "loving", "worries" and "cried" count. A sounds-like search for "Stephen" also lists the lines that say "Stephen" before a shorter one that says "Steven".
  • A speaker in your search keeps to what that speaker said. Typing "speaker Maya pricing" read the name as "Maya pricing", found no one by that name and searched everyone, so Bob's lines about pricing came back. With the name typed last it did find Maya, but kept every line of any recording she was in, Bob's included. It now reads Maya, searches what she said about pricing, and treats a typed speaker exactly as one picked from the list. The emotion and sounds-like filters also keep to the chosen speaker, where before they found nothing.
  • People outside the top of the People page keep their names. The same people were drawn from that shortened list everywhere else too, so they showed up as bare "Person 74" strangers: renaming one offered an empty box instead of the name they already had, the page you landed on kept the old label as though the rename had not saved, a recognized public figure lost their ⭐ banner, and "Move to existing…" would not offer them at all. They now read the same as anyone else.
  • Ask answers with the line that answers the question. In an interview, a meeting or a podcast the question is usually said out loud before its answer, and an answer composed without the local model led with that restated question ("Does anyone know when the launch is?") while the reply sat behind it or was left out. It now opens with the reply. Its supporting quotes stay on topic, so a pasta answer no longer adds a line about watering tomatoes; a correction made later in the same recording ("it moved to Thursday") is kept instead of being dropped as a repeat; and a citation points at the words it quotes rather than at another line of the same passage. An answer drawn from several files quotes the sentences that answer, not whole passages with their greetings and sponsor reads. And when the local model replies that your recordings don't mention something, Ask no longer shows that reply as a sourced answer.
  • One person, one face. Opening a person used to show every single detection — a talking head sampled through an 18-minute video filled the page with 237 crops of the same face. The strip now shows the best crop of each appearance — one per clip in the list beneath it — with Show every crop when you want the rest. Splitting still reaches every individual crop.
  • The same face no longer fills Faces & People as a dozen strangers. Indexing attached each new face to whoever it matched and invented a new person when nothing matched, but never merged the near-duplicates that creates — so a face caught at a hard angle became a second "Person 4312" and then collected more of the same. Recognized public figures are now de-duplicated by name, but that reaches only the faces the gallery could name: on one library the other 1,413 duplicate records were unnamed, which is why a sixth of the first page was still a duplicate of another tile on that same page. Indexing now folds those together as well, and Refresh faces → Merge duplicate people repairs a library that already has them. It only combines people who are clearly the same face — it never regroups anything else, and never merges two people you named differently.
  • Searching on-screen text finds the word you typed, not the letters inside other words. In a library full of slides, a search for "AI" put frames that say TRAINING, EMAIL or DETAILS ahead of the slide that says AI, and could push that slide off the first page; short searches such as "EV" or a ticker had the same problem. A word printed on screen now ranks above one that merely contains the letters, and an exact on-screen match is recognised as exact even when it also matches by meaning, so it no longer comes with the "closest guesses" note. Chinese and Japanese on-screen text is recognised as an exact match too, and a single character such as 猫 now finds the frames that show it. Searching your files' details also stops matching letters inside a folder name, a note or a chapter title: "war" no longer finds every file under a home folder named "edward" or a note that begins "Toward the end", and your account name no longer matches every file in your home folder.
  • A search with a filter in it searches the rest of what you typed. Once a search held a filter such as "videos", "4k" or "last week", the words around it lost their small words too: "videos with no agenda" looked for "agenda", a quoted "out of the office" became "out office", and "4k videos about the end of the world" looked for "end world". Only the words that led into the filter are dropped now, and a "no" or "without", or anything in quotes, stays as you typed it. Transcript search also stops treating two common words in a row as a quote: "going to launch" no longer puts "we are going to eat lunch" above the moment the launch was discussed.
Changed
  • "Move to existing…" is a search box now. Moving faces onto another person offers everyone in your library rather than a shortlist, so it is searched rather than scrolled: type a few letters and pick from the matches. It says how many matched, and will not let you move faces to nobody.
0.1.53
Added
  • See what a conversation library has learned about you. Settings → Memory now lists each library's facts — where you live, who you work for, what you like — with when you said it and in which conversation. Older values you have since changed are hidden by default and can be shown struck through, so you can see that "lives in Munich" was replaced by "lives in Berlin". Forget the conversation that said Berlin and the list updates on the spot: Munich is current again.
  • Keep your library working after an update. MediaFind checks that the library it already built still matches the models the new version uses. If something no longer matches, a banner explains what is affected and offers to refresh only what needs it.
Changed
  • A default index now covers every channel it can. Faces, object boxes and the shadow ("what's missing") scan used to need a second, separate pass that most people never made, so their channels searched an empty table. They now ride an ordinary index run: - Faces are detected and clustered on a normal mediafind index (and from Add Media, the API and a download), instead of only on --faces. Still Pro and still entirely on-device — a free install quietly indexes everything else and says so rather than failing, and asking for faces outright still shows the upgrade prompt. mediafind index --no-faces turns it off for one run and MEDIAFIND_FACES=0 turns the default off everywhere, the API included. - Objects are detected during indexing alongside logos and actions, wherever a grounded detector is installed. None ships with MediaFind, so on a stock install this costs nothing — no download, no slowdown. Where one *is* installed, indexing detects on 24 frames per file, spread evenly over its whole length, so a long video cannot spend more time on objects than on everything else put together; the whole-library Detect objects job you start yourself is still exhaustive and replaces those boxes with the full set. MEDIAFIND_OBJECTS=0 keeps objects out of indexing entirely, and MEDIAFIND_OBJECT_MAX_FRAMES_PER_FILE moves or lifts the cap. - A shadow scan is queued automatically after an indexing run, so the channel follows a library that just grew instead of waiting for a button. Pro, deduped across back-to-back runs, and skippable with MEDIAFIND_SHADOW_AUTOSCAN=0.
Fixed
  • The small print under a setting is small print again. The one-line notes beneath controls — in Settings, under the folder and page-URL fields in Add Media, and beside the model pickers — were being drawn at full body size in the main text colour, so each note was larger and louder than the label it belonged to. All 48 of them now read as the quiet grey notes they were meant to be, matching the ones already inside dialogs, in both light and dark themes.
  • Quitting MediaFind no longer reports a crash. After the AI assistant had answered a question, quitting the app (⌘Q, the Dock, or logging out) could end with macOS saying "MediaFind quit unexpectedly", even though everything had been saved and the quit itself was normal. The on-device model's graphics memory is now released as the app closes, instead of being left for the system to tear down on its way out. An answer the assistant is still writing when you quit is wrapped up first, in a fraction of a second, so quitting mid-answer is not the one case that still reported a crash.
  • A memory library that ends just after New Year no longer loses its date questions. Asked in early January about "March", a library of dated messages, photos and notes treated that as a March still to come and searched nothing. It now reads it as the March just gone, as anyone would. Measured on a year-long life-log: the share of questions the date narrows went from 45% to 75%, and the right memory reached the top 20 results for 67% of questions instead of 60%.
  • Downloading media indexed it differently from adding it. The download routes' face default had drifted from POST /api/index despite a comment promising they matched; both now resolve it the same way.
  • Object search stored nothing at all. Where a grounded object detector was installed, the object channel reported itself ready, loaded its weights and ran on every frame — and saved not one box, because the detector was being called in a way that failed on every single frame, and the failure was swallowed. Nothing could fill the channel by hand either: the whole-library scan existed as a job that nothing could start. Both halves are fixed, and that scan can now be started through POST /api/objects/backfill.
  • A file indexed by a different visual model now says what that cost. Logo and action detection can only read frames encoded by the same visual model they themselves use, so when a file's frames had come from another one, both were skipped in silence — the file simply ended up missing facets a matching model would have given it, with nothing to point at afterwards. The index run now warns, naming the file, the model its frames were encoded by, and which detection that cost.
  • Adding a YouTube video by its link works again. Pasting a YouTube URL into Add Media downloaded part of the video and then stopped with a "403 Forbidden" error. YouTube changed which of its players still hands out video, and the downloader MediaFind bundles had been left behind on the old one. Builds now require a downloader recent enough to ask a player that still answers. If you installed the optional download extra yourself, pip install -U yt-dlp picks the fix up.
  • An out-of-date downloader says so, instead of blaming the link. When the bundled YouTube downloader falls behind, YouTube answers it with "Sign in to confirm you're not a bot" — which MediaFind read as the *source* demanding a login, and reported an ordinary public video as login-only or DRM-protected. That is a permanent-sounding verdict about a link that was never the problem. The answer is now recognised for what it is, and the message names the stale downloader and how to update it.
  • The three Conversations cards in Add Media line up. *From a file*, *Paste a chat* and *A folder of notes* now sit side by side at the same height, each with its library field and Import button in a footer along the card's bottom edge, so the three buttons share one line however long a card's description runs. The line naming the file you chose, and the one reporting what an import brought in, are small and grey again rather than full-size body text.
0.1.52
Added
  • Search only your conversations, or just one. Find's scope button, left of the search box, now offers Conversations, and an open conversation has Find in this conversation, which searches that conversation alone. A date in the search still narrows it, and the × beside the scope button searches everything again. mediafind search --folder mf-text:// and the search API's folder limit a search the same way.
  • Manage your conversation libraries in Settings. A new Memory section lists each library with how many conversations and messages it holds and when it last changed. You can export a library to a file that imports again, here or on another computer, delete a single conversation, or reset a whole library, and a delete always asks first. It also shows how facts are read from your conversations: with the built-in patterns, the on-device model, or not at all.
  • Import a folder of notes from the app. Add Media → Conversations has a new card, A folder of notes: choose a folder of Markdown or text notes, such as an Obsidian vault, and each note becomes a conversation you can read, search and ask about. The desktop app opens a folder dialog; in a browser, paste the folder's full path. Hidden folders such as .obsidian and .trash are left out, a note is dated only by the date in its front matter, and importing the same folder again skips the notes that haven't changed.
  • Export a conversation library with personal information masked. Tick redact PII beside a library's Export in Settings → Memory for a copy you can share: emails, phone numbers, card and social security numbers in its messages and titles are masked, the people in it are called Speaker 1, Speaker 2 and so on, and file names and folders are left out. The library itself isn't changed, and the file's name ends in _redacted. Importing the copy adds its conversations as new ones, so it never takes the place of the conversations it was made from.
  • A larger on-device model for Macs with 32 GB of memory or more. The AI assistant model pickers and Settings now list XL, a larger on-device model. It is an 18.6 GB download and uses about 19 GB of memory while it runs, so it is offered only on a machine with 32 GB of memory or more and shown as unavailable on one with less. Nothing changes unless you choose it: the default and recommended models stay the same, and it downloads only after you pick it.
  • Ask answers from your conversations too. Answers on the Ask page can cite messages from your conversations next to moments in your recordings. A message shows who said it and when, and clicking it opens the conversation at that message. When only your conversations match and none of them answers the question, Ask says "Your conversations don't cover that." and shows the closest messages instead of guessing.
Changed
  • The on-device model reads more out of your conversations. When it reads facts (MEDIAFIND_MEMORY_FACTS=llm), it now catches corrections ("I meant 40, not 30") and how many of something you bought. Facts an earlier version read are replaced the next time a conversation is read, instead of staying beside the new ones.
  • A crash inside a transcription or visual-analysis model no longer closes the app. On computers with 16 GB of memory or more, indexing now runs those models in separate helper processes by default. When one of them crashes on a file, that file is listed as failed and can be retried, and the rest of the library keeps indexing. Computers with less memory keep the previous behaviour, because the helpers use about 2 GB more while indexing. MEDIAFIND_MODEL_WORKERS=0 turns them off and =1 turns them on.
Fixed
  • A long recording no longer gets stuck repeating one sentence. Transcribing a long file could get the recognizer stuck: partway through, it would stop following the audio and write the same phrase over and over to the end of the file — on one 18-minute interview, for the last 14 minutes. Those words were never said, so they showed up in search and in answers. Transcripts are now written without feeding each part of a file back into the next, which is what caused it. Where you have given MediaFind names and terms of your own, which it can only carry that way, a transcript that runs into a loop is written a second time without them. Files transcribed before this are not repaired automatically: re-transcribe an affected file from its page, or Settings → Transcription, to get a correct transcript.
  • An export no longer waits for indexing to finish. Every background job used to wait in one line for two workers, so a clip, an export or a model download you started while a big folder was indexing sat there until the indexing was done. One more worker is now kept for work you are waiting on. Indexing keeps the two it had, so it is no slower, and library-wide clean-up work waits behind what you started yourself.
  • Adding the same folder twice queues it once. Asking to index a folder again with the same settings while that request is still waiting to start no longer queues a second copy of the work.
  • A briefly busy library no longer fails indexing. When the library database was busy for a moment, the index job failed and had to be started again by hand. It now waits a few seconds and tries again, up to three times.
  • A job that hangs no longer holds up the rest. A job that stopped responding kept its worker until MediaFind restarted, and two of them stopped all background work. Once the two-hour limit marks such a job stuck, the other jobs carry on.
  • Jobs waiting when MediaFind closed are queued again, oldest first.
  • A paused indexing job stays paused when MediaFind closes. Pausing indexing and then quitting failed the job at the next start, with "MediaFind closed before this job finished". It now comes back paused, and Resume picks it up again. The same goes for a job you pressed Pause on that had not paused yet when you quit, which used to start again on its own.
  • The job history no longer grows forever. Finished jobs older than 30 days are removed, and only the 2,000 most recent completed ones are kept. Failures stay for the full 30 days so you can still see what went wrong.
  • A cancel during a file's visual analysis is noticed sooner. Indexing used to notice a cancel only once that file's frames had been analysed. It now notices within a couple of seconds and does not save the file.
  • The activity panel no longer piles up requests when MediaFind is busy. While indexing slowed the app down, the panel kept asking for job updates before the last answer arrived. It now waits for that answer.
  • Background jobs no longer keep files open after each check. Every look at a job, such as the activity panel's updates or a job reporting its progress, opened the job list's database and left it open until memory was next cleaned up, and indexing's check on each file did the same. While a job ran and the panel followed it, as many as 30 files stayed open that way, and queuing 200 jobs in a row left about 180 open, where a Mac app may have 256 open at once. Each check now closes its files as soon as it is done.
  • A damaged library no longer pushes your good backups out. MediaFind keeps your last seven backups. When the library file itself was damaged, each new backup, including the daily automatic one, copied the damage and deleted the oldest good backup, until none was left to restore. Each backup is now checked before it is kept: a damaged library is not backed up, the error says so, and every existing backup stays.
  • A search no longer comes back empty when the words were said. Searching a transcript could show nothing although the words are in it (about a third of our test searches did, one-word and full-sentence alike) when the passages that first looked close were then ruled out on a closer read. Search now shows where the words were said, as it already did when nothing looked close at all.
  • A new question in Ask is no longer mixed up with the one before. Once a conversation had a turn, any question starting with "where", "when" or "why" counted as a follow-up and was searched together with the earlier question's words, so asking "Where is the venue for Saturday?" after something unrelated could miss the passage that answers it. Such a question now stands on its own unless it points back, as "why?" or "when did they arrive?" do.
  • Asking the same question again gets the same answer. In an Ask conversation, a question asked a second time skipped what had already been quoted and could answer with an unrelated passage instead. Only a follow-up, such as "tell me more", now looks past what was already quoted.
  • A note dated by day shows that day. A note whose front matter gives a date without a time, such as date: 2025-06-02, showed a day early in the Americas and anywhere else behind UTC: in Conversations, in Find's results and in Ask's sources. It now shows the date as written, without a time. A date with a time still shows in your time zone.
  • Face detection's runtime can no longer report usage to Microsoft. The runtime behind face detection comes with usage reporting switched on, keeping a device ID and a queue of events to send to Microsoft. MediaFind now switches it off before the runtime starts, so turning on face detection keeps the promise that nothing leaves your computer.
  • Keyboard users can reach what used to need a mouse. Space now ticks a search result's select box (and the ones on face and speaker clips) instead of opening the result. Choosing people to merge, Insights' month bars and calendar days, and a file's context menu (Shift+F10, then the arrow keys) work from the keyboard; the closed details drawer no longer hides Tab stops off-screen; and a transcript line's play button shows a focus ring.
  • Focus stays where you were after a keyboard action. Saving or cancelling a transcript edit or a speaker rename, closing a file's context menu or Create's shortcuts panel, and renaming or opening a Create project used to drop focus to the top of the page. The shortcuts panel also keeps Tab inside it, and the backup list in Operations no longer takes focus away every few seconds.
  • Scope and channel pickers close when you Tab away from them, and the activity dock's minimize button works with Enter and Space.
  • The quarantine confirmation shows a readable plan. Before moving a low-quality file, or the extra copies in a duplicate group, to quarantine, MediaFind asks you to confirm. That dialog printed its markup as text, so the files that move, the destination folder and the note that you can restore them were buried in tags such as <p>Moving <b>….
  • Quitting or crashing during a backup or a restore no longer leaves a hidden copy of your library behind. Backups and restores work on a hidden copy and tidy it away when they finish, so one interrupted before then left a file as large as your library on disk for good. On Mac and Linux the next backup or restore now removes it.
  • An interrupted restore can no longer leave a library that is half old, half backup. When another MediaFind process had the library open, a restore cut short between two of its last steps could combine the two without any error. That window is closed, and a restore now refuses with a clear message while another process is still using the library.
  • A crash while moving files to or from Quarantine no longer hides them. A move cut short could leave a file in neither the library nor the Quarantine list, or take it off the list without reaching the Trash. Each move is now recorded before it is made, and MediaFind settles an interrupted move at its next start, as soon as it can tell where the file ended up. A permanent delete that was cut short no longer leaves the file searchable.
  • An interrupted index run no longer leaves files half-indexed for good. If MediaFind was killed or ran out of disk space while indexing a file, later runs treated the file as done and never added what was missing, such as an image's visual search data or a recording's summary and chapters. Such a file is now indexed again on the next run.
  • A library file that became empty no longer opens as a new, empty library. When a good backup exists, MediaFind now stops at startup and names the backup to restore.
  • A damaged job list or usage record no longer stops MediaFind from starting. The damaged file is kept aside, the log says where it went, and a fresh one is started.
  • The list of files that failed to index and your setup progress survive a crash while saving. An interruption at the wrong moment used to reset them.
  • A disk that fills up for a moment during indexing no longer leaves a file incomplete for good. When saving part of a file failed for lack of space, indexing carried on without it, so the file could be stored without its summary and chapters, categories, meeting notes, music detection, a thumbnail, a face crop or the actions found in a video, and every later run skipped it as already indexed. The file is now listed among those that couldn't be indexed, with the disk-full error, and Retry, or indexing the folder again once there is space, indexes it in full.
  • "Add to Claude Desktop" no longer wipes out a config file it can't read. When Claude Desktop's config file had a comment or a trailing comma in it, adding MediaFind rewrote the file with MediaFind alone, and a second click overwrote the backup too, losing your other servers and settings. MediaFind now leaves such a file untouched and shows what to paste into it by hand. It also stops before adding a server that can't start on this install, and when it can't write to Claude Desktop's folder it says so plainly instead of showing an error code.
  • The update channel you pick stays picked. Switching Settings to the beta channel lasted only until MediaFind quit.
  • The demo's example clips stay gone when you restore a backup. Turning demo mode off removes the example clips it added, but restoring a backup, or importing a library export, made while demo mode was on brought them back as files that could not be found. They are now left out; an example clip you have added again since is kept.
  • Detecting logos says so when it can't run. Where logo detection isn't available, it reported "Detecting logos complete" after doing nothing; it now says it was skipped and why.
  • A job you cancel no longer looks like it failed in the activity dock, and finished jobs can be dismissed from the Operations list like failed ones.
  • MediaFind fits a phone-sized window. At 360 pixels wide, pages no longer scroll sideways once the Back button appears or when the Video Editor lists a long file name, a notification no longer covers the top bar's buttons, Operations shows backup and job names instead of squeezing them to a letter, and Create's timeline, audio and captions buttons wrap instead of running off the screen.
  • At 200% zoom or in a short window, you can see Ask's answer and reach the whole sidebar. The question box no longer stays pinned over the answer, and every sidebar entry can be scrolled to.
  • Small controls are easier to tap: a voice's select, rename and link buttons (rename and link now show on touch screens, which have no hover), Settings' checkboxes, and Create's back, project settings and shortcuts buttons.
  • You can import a chat saved as a Markdown file. Add Media → Conversations now offers .md files, which the importer already understood.
  • "Suggested voice matches" says it is part of Pro on a free install, instead of showing the server's error message as plain text.
  • Searching for a word no longer turns up other words that contain it. When search shows where your words were said, "race" also found "traced", "yard" found "vineyard" and "cross" found "across". A match now has to start a word, so "child" still finds "children" and "launch" finds "launched". Words outside the plain English alphabet still match anywhere: Chinese and Japanese are written without spaces, so a word inside a sentence has no start to find, and an accented word such as "café" follows the same rule.
  • URL downloads refuse addresses on carrier-grade NAT and Tailscale networks. A pasted link to a 100.64.x.x address, the range carrier-grade NAT and Tailscale use for devices on a private network, was fetched, even after a redirect, although local and private addresses were already refused.
  • Create's shortcuts leave text fields and hidden rooms alone. ⌘Z while typing in the Bin search undid your last timeline edit instead of the typing, ⌘D duplicated a clip, and Backspace in the Deliver room's preset menu deleted the selected clip, which that room doesn't show.
  • Timeline edits in Create do what they say. Dragging a clip you hadn't selected now moves it. "Other track" no longer drops a video clip onto the audio lane. A refused edit shows the reason instead of "move 422". Split moves the timecode with the playhead, and "Removed 2 clips." stays on screen.
  • People and Meetings on the free plan say they are part of Pro instead of "No people detected yet" or "recordings ready", which wasn't true for a library with people or processed meetings from a trial.
  • A long collection or saved-search name no longer pushes Organize's buttons off-screen.
  • Create and transcript export give a clear refusal instead of a server error for a file that isn't in your library, a folder path that isn't absolute, or a snapshot name with a control character or hundreds of characters.
0.1.51
Added
  • Memory answers can see your facts, as an experiment. Memory search and answers can now show the facts your conversations established next to the messages they came from (--include-facts 1 on mediafind memory search and ask, include_facts in the memory API, the MCP tool and the OpenClaw backend), so an answer can tell that something you said later replaced what you said before. It is off by default because it does not make answers better yet: in our tests it fixed some questions about things that had changed, but broke more than it fixed, most often questions that draw on several conversations at once.
  • Conversations keep their facts up to date on their own. What you said about yourself (where you live, what you own, how many of something you have) used to be read only when you asked with mediafind memory facts --extract. Now every conversation you save or import has its facts read for you: in the app that happens in the background, so searching or saving never waits for it, and a one-off mediafind memory ingest reads them before it exits. The on-device AI model can do the reading instead of the built-in patterns: it catches what people mention in passing ("I just got a new guitar, how do I set it up?"). Set MEDIAFIND_MEMORY_FACTS=llm, or pass --extractor llm. Nothing leaves your computer either way.
  • Forget one conversation, not just all of them. You could clear a whole memory library or nothing; now mediafind memory forget <library> <conversation> removes a single one, and everything that hung off it goes too — its turns, anything the assistant had quoted from it, and the facts it had established. If that conversation was where you said you had moved, the earlier answer it had replaced becomes current again rather than leaving a blank.
Fixed
  • "Add to Claude Desktop" finds Claude Desktop on Windows and Linux. It looked for Claude Desktop's settings where the Mac app keeps them, so on Windows and Linux it always said the settings folder was not found and asked you to paste the settings by hand, even with Claude Desktop installed.
  • Ask works on the free plan again. With the evidence setting left at its default, "spoken + visual", every question came back as "People & Faces is a Pro feature", because visual evidence is part of Pro. Free installs now start on spoken evidence, and choosing visual evidence opens the upgrade page instead of an error.
  • The free plan no longer shows who appears in a file, or when a brand does. After a Pro trial ended, a file's details still named the people in it and listed when each brand was on screen, while everywhere else that information needed Pro. Those lists now stay empty until Pro is unlocked.
  • Typing a person's name into search now needs Pro, like the person filter. On the free plan, a search such as "person Alex" still narrowed the results to where that face appears, using face data kept from a trial. It now asks you to upgrade, and mediafind search and the MCP search tool refuse it too.
  • Saved searches stop running Pro searches on the free plan. A search saved during a trial with a person filter, or with the People or brand-logo channels, still refreshed with them after the trial ended. Saving or refreshing one now needs Pro, and a refused save or refresh leaves the saved search as it was.
  • Clicking a collection in a file's details on the free plan opens the upgrade page instead of doing nothing. Collections made during a trial still show there and in search. "Delegate as a task" also stops saying "Understanding your request…" once it has opened the upgrade page.
  • Importing a conversation shows its result again. A pasted chat or an imported export was saved, but the page reported that the import had failed — no summary, no link to open it, and a file with the same name was never offered as a replacement.
  • Sending search results to Create no longer fails because one file moved. If any picked file had been moved or deleted since it was indexed, the whole project failed with the message "[object Object]"; that file is now skipped and the rest are placed.
  • The app's scripts are always sent as JavaScript. On a Windows computer whose settings label .js files as plain text, the app's window would have loaded with every button dead.
  • A transcribed file no longer says it has no transcript. Without a local AI model the file view (and mediafind file) read "no summary — this file had no transcript" even for a fully transcribed recording; it now says "no summary yet".
  • Ask's follow-up suggestions stopped offering nonsense topics. Suggestions such as "Tell me more about deeply" came from words the sources mention only once; they now use only terms the sources repeat.
  • A channel restriction now applies to filter-only searches too. Asking for "videos" while restricted to transcripts returned videos with no transcript; browsing lists files and runs no channel, so it now answers nothing rather than claiming a match. Unrestricted browsing is unchanged.
  • CLI search now understands the same queries as the app. mediafind search uses the shared search service, so queries such as "videos from last week about the budget" search for "budget" within the matching kind and date filters. Filtered searches also use the shared wider candidate pool. Existing search flags and output formatting are preserved.
  • A person or speaker filter now applies to filter-only searches. Asking for "videos" (or "4k videos", "no transcript" — anything the parser claims entirely as filters) browsed the library and dropped the filter, answering with files the person never appears in. Searches with words in them were never affected.
  • Importing a big export no longer freezes the rest of that library. The import now runs in the background with the activity dock tracking it, and it steps aside between conversations instead of holding the library for the whole run. Reading or searching the same library while an import runs used to time out; now it waits for one conversation at most. Clearing a library is refused while one of its imports is still running, so it cannot be left half-imported.
  • The MCP server starts again. pip install "mediafind[mcp]" had begun pulling in mcp 2.0, which renamed the class the server is built on, so mediafind-mcp and the app's --mcp mode stopped with "needs the 'mcp' package" even though it was installed — including in the 0.1.50 builds made from that install. The extra now pins the supported line, and when a newer mcp is present the message says so instead of asking you to reinstall it.
  • Quitting while MediaFind indexes no longer fails the indexing. Closing the app mid-index used to hold the quit while indexing ran on, then list the job as failed with "server restarted during job", as if MediaFind had crashed. Indexing now stops when you quit, waiting a few seconds at most for a file in progress, and carries on the next time MediaFind opens, skipping the files it had already finished. Other work that quitting cuts short, including a paused index, now says MediaFind closed before it finished.
  • Stopping mediafind serve no longer counts the files it was indexing as crashes. Its exit skipped the server's own cleanup, so every file in flight was charged a crash, and after three stops a healthy file was skipped as if it were broken. The same exit printed a warning about a leaked semaphore.
  • Connecting an AI agent works on Windows and Linux. "Connect to AI agents" wrote a configuration that started MediaFind with an option the Windows and Linux app did not accept, so the agent saw a server that never answered. It now starts the MCP server there too, as it already did on macOS.
  • Refreshing embeddings now covers the text in images and video frames. After a search-language change, the refresh updated transcripts and summaries but left on-screen text behind, so searching by meaning for what a slide or sign says stopped finding it (exact words still matched) while the app reported nothing left to refresh. On-screen text is now refreshed along with the rest.
Changed
  • Internal tidying of how the app's parts are layered. Nothing changes for you.
0.1.50
Added
  • Your notes can become memory, not just your chats. Point MediaFind at a folder of Markdown or text files and it reads the whole tree as conversation-style memory you can search and ask questions of: mediafind memory ingest ~/Notes --library personal --format note. Long notes are split into passages, so a phrase buried on the last line of a long note is still findable — storing a note whole would quietly lose its tail. A note's title comes from its front matter, its first heading, or its filename.
  • Conversations remember what you told them, and notice when it changes. Until now a conversation library could find the sentence where you said something; it had no idea the sentence was still true. MediaFind now reads plain statements out of your own turns — where you live, what you do, who you work for, your pets, your preferences, what time you usually do things — and keeps each one with the turn that said it and the date it was said. When a later turn changes it ("I moved to Austin") the old fact is marked superseded, and when a later turn takes it back ("I don't live in Austin anymore") it is marked retracted. Nothing is deleted, so you can still see what was true when. Two statements made at the very same time both stay, listed as a conflict rather than one silently winning.
  • Memory search can bound a question's period at both ends. Asking "what did I photograph in mid-January" searched January onwards — the end of the period was never a bound, because in a conversation an event is usually recounted *after* it happens. That does not hold for a library of timestamped things (messages, photos, calendar entries), which cannot postdate the period they belong to. The new question_span="span" bounds both ends for those libraries and understands "early/mid/late <month>" and a bare "March 14", which previously narrowed nothing. Off by default, so conversation histories are unaffected. On libraries like those it finds markedly more of what you asked for, and searches faster.
  • A bundled gallery of well-known voices. MediaFind now recognises a notable voice with nothing to set up: it ships with reference voiceprints for public figures, built from freely licensed recordings that are credited inside the app. Matches stay provisional — ⭐ in the voice roster, ✎ to confirm, ✕ to dismiss.
  • A much stronger voice model. Diarization and voice recognition now run on a new voice model that separates similar voices about twice as clearly as the old one, and it is what makes recognising a known voice reliable. It arrives as a small one-time download; until it does — and on any machine that cannot fetch it — MediaFind keeps using the previous model, so nothing stops working. Voiceprints taken by the old model cannot be compared with the new ones, so recordings you have already indexed keep the voices they have until you re-detect them (the Speakers panel offers this).
  • Recognition by voice print — the voice counterpart of face recognition. Every voice MediaFind separates out is matched, on this machine, against the voices it can already put a name to: voices this library knows (one you named, or one linked to a person whose face was recognized), reference clips you add yourself, and the bundled gallery of well-known voices — so a name given once reaches every podcast, voice-over and audio-only file that voice speaks in. Matches are provisional (⭐ in the voice roster; ✎ confirms, ✕ dismisses and is remembered), and run after each diarized file and on demand via Identify notable voices, mediafind speakers --identify or POST /api/speakers/identify.
  • Memory search can follow who said something. When a question names someone who speaks in your conversations ("what did Maya say about the telescope"), MediaFind can also search that person's own messages, instead of only the messages that mention their name by the way. Turn it on with --speaker 1 on mediafind memory search and mediafind memory ask, or speaker: 1 in the memory API; it is off by default. Names match as written and as whole words, so "Will" the person is not "will" the verb.
  • Transcripts and meeting notes export as Word documents and PDFs. Pick *Word (.docx)* or *PDF* next to the Download button on any file, use the new Export buttons on a meeting's brief, or export a live meeting from the panel while it is still running — the meeting document leads with the overview, action items, decisions and key points, then the full transcript with speakers and timestamps. Both files are written on this machine with nothing uploaded, and Chinese, Japanese and Korean meetings come out as readable text rather than empty boxes. From the command line: mediafind export FILE --format docx --out notes.docx. Word needs nothing extra; PDF needs pip install -e '.[docs]' on a source install and is built into the app.
  • Live meetings can record both sides of a call on a Mac. A new audio source, *Call on this Mac*, records what the Mac plays (the other people on a Zoom, Teams or Meet call, in the app or in a browser) together with its microphone, so a call on headphones is transcribed in full. It needs macOS 13 or later, and macOS asks once for the Screen & System Audio Recording permission; MediaFind keeps only the audio. In the CLI it is mediafind meetings live --source call. Pro, like live meetings.
  • Transcription learns your names and terms. Settings → *Names & terms* takes people, products and jargon, one per line, and every file and live meeting is transcribed with them in mind, so a colleague or a product comes out spelled as entered instead of as a sound-alike. About fifty short terms fit and the first ones win; a change reaches the next utterance of a running live meeting, the live panel says how many terms are in force, and MEDIAFIND_ASR_VOCABULARY sets the list for the command line. A recording's transcript glossary still fixes what was transcribed before.
  • A Turbo transcription tier — the most accurate an 8 GB Mac can run. The transcription picker gains *Turbo*: close to the Large tier's accuracy at the Medium tier's memory footprint, and quicker to transcribe than either. Large itself wants 16 GB of RAM, which has kept it out of reach on an 8 GB Mac; Turbo is the first tier to bring that level of accuracy — on accents and non-English speech especially — to those machines, and it keeps word-level timestamps on Apple GPUs.
  • The on-device answer models step up a generation. The bundled Mini tier and the downloadable Small, Medium and Large tiers all move to newer models, staying in the same size class so the machines that ran them still run them; the app download grows by about 130 MB. A tier you already downloaded keeps working offline — the model it shipped with before stands in until the new one is fetched, so an update never turns a working setup into a multi-gigabyte download in the middle of an Ask. The largest tier deliberately stays where it is: its next generation is big enough that only a 32 GB machine could run it.
  • MediaFind can be the memory behind an OpenClaw agent. mediafind memory openclaw-backend serves a conversations library to OpenClaw's simple-memory plugin: every turn the agent sees is saved, the turns that match a question are handed back with their date and speaker, and a reset forgets everything. Local and keyless: the memory side spends no model tokens.
  • Faces too blurry to identify are no longer shown. A face can be large and still be a smear — motion blur, defocus, a compressed face in the background — and a size check alone could never catch it, so People filled up with faces nobody could put a name to. MediaFind now judges how sharp each face actually is and leaves the hopeless ones out of People, the file view, appearances and search results, while the live detector stops storing them at all. A person whose every shot is that soft drops out of the People grid, and most of those turn out to be junk rather than people. Nothing is deleted — hidden faces still group together and still carry identity — and a person you have named or confirmed as a public figure is never hidden. Existing libraries are judged once in the background, with no video re-processed. Turn it off with Settings → Hide blurry faces; tune the floor with MEDIAFIND_FACE_MIN_SHARPNESS.
Fixed
  • Architecture guards now include async handlers and all six mutable model settings, with router prefixes respected when checking duplicate routes.
  • A busy library is no longer reported as corrupt. Startup continues after the health check, and the status indicator explains when another MediaFind process is using the library.
  • Reinstalling an older build detects a newer library before making a backup. Startup names the unknown revision and explains how to reinstall the newer version or restore a compatible backup, without creating a copy on every launch.
  • Closing the desktop window requests a graceful shutdown. The server closes its job queue and index pool, with a bounded wait so quitting cannot hang.
  • Expensive and state-changing GET routes reject cross-site requests, including library ZIP exports, waveform generation, lazy cache writes and library analysis.
  • MediaCreate validates client media paths before probing or saving refs. Every route that takes a media path now checks it — adding a clip, music, handoff, compose, folder import and importing a .mfcreate document. A path that is only recorded is checked by extension, so a project whose source drive is unplugged still opens; a path that is about to be probed must also exist, and adding a clip needs an indexed source or an existing project ref (which keeps imports that are still being indexed working).
  • Downloader destinations must already exist. A request can no longer make the app create a folder tree anywhere it can write; both download routes reject a missing folder before enqueueing and the worker repeats the check. Only the default downloads folder is still created automatically, and mediafind download --out ./rel keeps working.
  • A file whose name looks like HTML is shown as text. Three places rendered a path straight into the page — the timeline export's status line, the video tools' "saving to" line and the relink candidate list — so a file named like a tag reached the UI unescaped. Library data now goes through one shared escaping primitive.
  • Transcript glossary rules no longer grow transcripts on every apply. An expansion rule such as "Ann" → "Ann Smith" was re-applied whenever another rule was added, on Apply, and on every re-index, turning the text into "Ann Smith Smith Smith…" in place. Applying a rule twice is now a no-op, and a replacement containing a backslash is inserted as typed instead of being read as a pattern (which could break indexing until the rule was removed).
  • Free accounts no longer receive People and Meetings data on the side. The status payload the home page loads, and the "when does X appear" events query, handed out names, face crops and meeting summaries that the People and Meetings pages themselves refuse without Pro.
  • Demo mode now blocks every way of adding conversations, not only the file import: the structured ingest and append endpoints, the MCP memory tool and the command line refuse the same way.
  • Asking the assistant to find duplicates no longer piles up copies of every duplicate group in Cleanup on each run; its scan replaces the last like a Cleanup scan does.
  • Ask answers that combine spoken and visual evidence are no longer always marked "weak match" and no longer always re-search once: their evidence kept a rank-fusion weight where the app expected a relevance score.
  • Two unrelated noisy recordings are no longer reported as sharing a music track, and a colour search limited to a folder no longer comes back empty when other folders have stronger matches. Object-box backfills no longer lose every stored box when a file moved during the run, and the MEDIAFIND_SCENES / MEDIAFIND_EMOTIONS switches now do what the docs say.
  • The command line indexes folders the app can open. mediafind index . stored relative file paths that only resolved from that directory; the app could not play them, export called them missing, and adding the same folder from the app indexed everything a second time. Targets are stored as absolute paths now. A file the indexer gave up on is reported as failed instead of "indexed: nothing", a run in which every file failed exits with an error, and mediafind people --forget shows what it would delete and asks for --yes before wiping faces, names and voice links.
  • Smaller: evidence coverage receipts now count only the files in the asked scope, and an empty scope no longer reports visual coverage as complete; two different files with the same name get their own audiograms; the MCP list_people tool honours MEDIAFIND_MCP_ALLOW_CONTENT=0.
  • Free accounts could see face data through the speaker link suggestions. The suggestions list and its confirm/dismiss actions now need Pro, like the People and Identities pages they draw from.
  • Downloading a second video with the same title kept the first file. The download was skipped by the fetcher, reported as done, and the first file was re-labelled with the second URL. The second video now gets its own name, and a download that produced no new file fails instead of succeeding.
  • Indexing an external drive or a NAS share picked up macOS metadata. The hidden ._name companions and the drive's .Trashes folder were indexed (a deleted file became searchable), and each companion failed again on every watch tick. They are skipped now.
  • A folder added twice under two spellings was indexed twice. Reaching a library through a symlink, or typing its path with different capitalisation, created a second copy of every recording and every search hit. A second spelling of an indexed file is now recognised and skipped.
  • Redacted speaker and clean-transcript exports printed speaker names. They now use "Speaker 1", "Speaker 2" like a redacted share.
  • Conversations could not be exported from the command line, and could be sent to an NLE timeline or to Resolve as a phantom clip. The CLI exports them as text now, and timeline exports skip them with a warning.
  • A mixed-frame-rate EDL placed 29.97 fps clips 1.2 s off after 20 minutes in editors that honour the file's frame-count mode. Source timecodes now follow the mode the file declares.
  • Merging a voice into one detected by the older method re-flagged the good recording. Re-detect no longer overwrites a recording's real voice-model stamp or re-runs on recordings the current model already handled.
  • "The transcription model is still downloading" could be shown forever. After a failed or cancelled download the app kept claiming a download was in flight, hid the re-transcribe banner and refused to re-transcribe. Missing weights are now reported as missing, with a hint to download them.
  • Model pickers ignored an environment-pinned tier. With the transcription or AI tier pinned by an environment variable, choosing another tier downloaded the pinned one and reported success. Pinned tiers are now shown locked, and the app says when a different tier was used.
  • A conversation ingested from the command line could not be re-imported or replaced from the app. It now offers the usual Replace choice.
  • Light mode no longer flashes dark on the way in. The theme was applied by the last script on a 1.5 MB page, so a light-mode viewer watched the dark default paint first; it is now set while the page's head parses.
  • Smaller fixes: vendor ffmpeg builds (n7.1, git and dated snapshots) were called too old; two simultaneous downloads with the same name could overwrite each other; a selection package export with missing files now says which are missing; CSV exports neutralise spreadsheet formulas; VTT cue text is escaped; detaching a voice from an identity clears its link everywhere and lets it be suggested again; voice counts no longer include merged-away labels; the bundled Base model is not badged as downloaded for the Apple GPU engine; the local-LLM check follows MEDIAFIND_OLLAMA_URL; Memory no longer says "no evidence" when the only matching memory is longer than the answer budget; imports report conversations that held no messages; paste errors no longer name an internal file.
Changed
  • Importing a conversation looks like the rest of Add Media. Conversations is now two module cards in the same two-up grid as From local and From internet, each with the same icon, title and description: From a file, whose styled Choose file… button replaces the browser's own file widget and names the file you picked, and Paste a chat. Each card imports only its own source into its own library field, so a leftover file can no longer win over a paste, and each reports its own result.
  • The Pro key is called a "subscription key" in the app too. The checkout email, the success page and the privacy page already said "subscription key", while Settings → *MediaFind Pro*, the MCP extension's config field and the API reference said "license key". The environment variable, the route path and the JSON field keep their names — they are wire contracts.
Security
  • Insights and visual evidence now enforce the People entitlement boundary, protecting stored person names, celebrity matches, face crops and brand signals. The Home dashboard still answers 200 for everyone with those sections emptied, so the timeline, duration histogram and free facets keep rendering.
  • Event questions no longer answer from retained Pro detections. "When does <person> appear?" was still served from prepared visual evidence, and brand appearances from stored logo rows, after entitlement lapsed.
  • A retained face link no longer names an unidentified voice. Voice recognition drew names from linked People rows without checking entitlement; a name you typed yourself is still yours.
  • Visual Ask answers and conversation history require Pro, including combined evidence and spoken follow-ups to earlier visual answers. Free memories omit person cards, speaker rosters hide face links, and direct face crops require Pro.
  • Retained workflow and background-job outputs respect Pro entitlements; free automatic reframing no longer consumes stored face detections.
  • Every composed API route is classified by an entitlement coverage test: ungated routes need an explicit free-by-design reason, and removed routes cannot leave stale allowlist entries behind.
0.1.49
Added
  • Conversations in the app. Import a ChatGPT or Claude export, a text file, or a pasted chat from *Add Media → Conversations*. The new Conversations page lists your libraries and their conversations and lets you read them. Ask a question about a library and the answer lists the messages it looked at; when the conversations don't hold the answer, it says so instead of guessing. Conversation results in Find now show who said what and when, and open the conversation at that message instead of a player. What an AI assistant saves through the memory tools shows up here too.
  • Re-detect speakers on recordings the older voice method split. Recordings diarized while the voice model couldn't load (fixed in 0.1.48) came out as one speaker, and their voices never match new recordings. When a library has them, the Speakers list in Faces & People now says how many and offers *Re-detect speakers*: free, it re-runs speaker detection on just those recordings in the background (one that changed on disk since it was indexed waits for its next re-index), and transcripts stay as they are. Voices from the older method are marked. Named ones stay, and unnamed ones that no recording uses any more are hidden.
  • Search understands more ways of naming a date. Searches now read "in May 2022", "May of 2022", "October 13, 2022", "from March 2020 to June 2021", "3 weeks ago" and "since April 2023" as dates. Before, a year after a month was ignored, so "in March 2022" was read as March of this year. Memory questions given as_of read the same dates and skip conversations from before them, except that "3 weeks ago" sets no window there: people often mention such things long after they happened.
  • Memory search can favour more recent turns and mark superseded ones [earlier]. Enable with recency=1 or --recency 1; this is off by default.
  • Memory search and answers can include surrounding turns. Use context=1 or --context 1 to include the turns just before and after each memory. This is opt-in and off by default.
Changed
  • Memory answers can combine, count and work out dates. The local model is now told the answer may need details from several memories, a count of them, or date arithmetic. It is still told not to guess, and to reply "I don't know" when the memories neither state nor imply the answer. A reply that opens with "I don't know." and then only says what the memories leave out now counts as an abstention.
  • The automatic update check is off until you turn it on. The box for it in Settings starts unticked, and nothing contacts GitHub until you tick it (or set MEDIAFIND_UPDATE_CHECK=1). Before, a fresh install checked for updates 3 s after every launch and showed the box ticked, and opening Settings could run a check by itself. "Check for updates" and switching channels still check on demand, and with MEDIAFIND_NO_UPDATE set, the button now says checks are turned off instead of "You are up to date". Installs that never ticked the box stop checking on their own; in the Mac app that is every install, since the setting did not survive a launch before this release (see Fixed), so tick the box again if you want the check.
  • Exports into a folder you chose never overwrite a file. The Video Editor and clip exports used to replace a file with the same name without asking, and different exports could share one name (GIFs of different ranges, frames at different times). A clash now gets name (2).ext, and frame, GIF and trimmed names carry their time or range. The default clips folder still overwrites same-name clips, so re-exported search clips don't pile up.
  • Restore refuses to run while jobs are running, and brings the restored library up to date. Restoring a backup from the app while an index, import, re-embed or other job is at work is refused with 409 jobs_running, naming the job types; only a model download doesn't block it, and a cancelled job whose work is still finishing counts as running. Restore then runs migrations and the index setup on the restored file, so it takes about as long as a start; if that fails (a backup from a newer version, a lock, a full disk) the previous library is put back and restore answers 400 upgrade_failed. A migration that can't get the database lock now waits 30 s and then says another MediaFind process is using the library, instead of swapping the file under it. mediafind restore on the command line still restores at once and leaves the upgrade to the next start.
  • AI clients see how a search was interpreted and can turn it off. /api/v1/search and the MCP search_media tool applied the query parser's filters without saying so: "landscape of competitors" became an orientation filter and found nothing for a verbatim quote. Responses now carry semantic_query, filters and filter_warnings, and parse=false searches the words as written.
  • MCP get_transcript returns one page at a time. A two-hour transcript was about 97k tokens, past AI clients' output caps with no way to page. It now takes start, end and max_chars (default 40,000, about 12k tokens), returns truncated and next_start, and no longer includes the joined text: a client reads segments and follows next_start, which is rounded down to hundredths so an echoed value can't skip a group of segments. Speaker names now come along with MCP search hits, and with transcript segments on MCP, /api/v1 and the SDK; MCP withholds them when content sharing is off.
  • Quoted names in a search are exact filters. collection "Client Interviews" and tag "…" now filter by that name; before, a multi-word name was read as its first word, and the quoted form stayed in the search text. A name that matches nothing says so in filter_warnings. A saved search using that form may change membership on its next refresh.
  • Every response except static files is marked Cache-Control: no-store. Now that the desktop window keeps its data store (see Fixed), the browser engine would otherwise write the page, JSON and generated images to disk, where they outlive a deleted library. /static sends no-cache, so an update can't pair new markup with an old stylesheet. The one thing the window still keeps outside data/ is its UI state in ~/Library/WebKit/MediaFind (macOS app), named in the privacy doc's "Deleting data" section.
  • Large libraries rebuild their vector-search graph once after upgrading. The graph's freshness signature changed (see Fixed), so the first search on a library above 10,000 segments builds it again. Each segment insert also pays a small bookkeeping cost, about 3 µs a row.
Fixed
  • Following the backup-recovery hint no longer deletes your good backups — when the library database was damaged, the startup hint named the newest backup without checking it, mediafind restore restored whatever it was given, and each restore first saved the damaged live database into the same seven-slot rotation, so every round pushed out the oldest good backup. The hint now names the newest backup that passes an integrity check, restore refuses a damaged backup, and a damaged live database is set aside as backups/<name>.corrupt-<time>…, outside the rotation. A healthy database that another MediaFind process is writing to is reported as in use, not damaged. Restoring a backup made by an older version also no longer leaves the default search failing until a restart (no such table: index_meta), summary search no longer answers from the library a restore replaced, and a restore during an index job no longer loses the job's writes (see Changed).
  • "Quarantine all" (Pro) no longer moves away the only copy of a file indexed under two spellings of its folder — a file reached as both ~/movies and ~/Movies on a case-insensitive disk, or under two Unicode spellings of the same name, was grouped with itself by the exact-duplicate scan, and the guard that keeps one copy compared the two paths as text, so "Quarantine all" moved the file out and reported its size as freed. The scan, the keeper guard and the quarantine step now compare the physical file, for the free tier's per-group quarantine too.
  • Redacted shares and exports no longer reveal the source file's name or path — a redacted share bundle's manifest.json still carried the full source path in its notes; redacted transcript exports used the file name as their title, Obsidian exports wrote the absolute path and file:// links, the saved file and the Video Editor's browser download were named after the source, and the CLI did the same. Call recorders name files by phone number. Redacted output now has a neutral redacted-source title, no source path or links, notes rebuilt from an allowlist, user tags redacted too, and files named redacted-transcript.<ext> (still never overwriting). A phone number next to another number, and numbers read out back to back, are now redacted, and E.164 numbers count.
  • "Re-detect faces" and licence activation no longer wipe every person name and voice link (Pro) — Re-detect cleared the faces before the re-grouping had taken the snapshot of names it carries across, and activating Pro ran the same detection on every file, so a trial user who had named family members lost every name the moment they paid. Names and links are now snapshotted before the clear, the re-grouping from that snapshot runs even if you cancel midway (the files done before a cancel used to lose theirs), and activation only detects faces in files that still need them.
  • More people (Pro) and speaker fixes — once the voice model became available, files diarized with the fallback were re-diarized and every named voice in them detached (the new voice is now mapped back onto the old label by talk time); merging speakers left action items on a label that could later be reissued to a stranger (freed labels are never reissued, and a library from before this release gets the same floor the first time it is read); two people you named the same kept the name on only one after Re-group; cross-file speaker matching depended on label order; a folder-scoped people search cut to its top 10 before scoping; voice links on people whose faces are under 24 px still transplanted; a looser Re-group handed a merged identity to the wrong person; merging same-named sources into an unnamed survivor dropped the name; and two voice-linked people swapping "Person N" numbers left both voices on one person. Re-group now also lets a named person whose faces are mostly inside a merged cluster keep its name and voice link over a nearer unnamed one.
  • The Mac app keeps its settings between launches — the desktop window's browser engine started in private mode, which wiped its storage as the window was built, so the theme, dismissed banners, the tour, the Asset Library view and the update opt-in were reset every time. They persist now. Settings still reset when port 7860 is taken at launch, because the fallback port changes the page's origin.
  • Search on large libraries no longer serves results from before a re-index or a model switch — on libraries above 10,000 segments, search uses a nearest-neighbour graph whose freshness check looked at the row count, the highest row id and the model stamps. Re-indexing a file with as many segments as before reused the freed ids, so the graph looked current: a search for the new content missed it and a search for the deleted content still found it, after a restart and in a second process with a warm graph (the MCP server beside the app). Switching between two models of the same width (English and multilingual text models) kept answering from the old model's graph, and mixed model stamps rebuilt the whole graph on every search. A change counter kept by the database now moves on every segment insert, delete and text or vector change, from any process, and the loaded model is part of the signature; merging speakers doesn't count as a change. The findings inbox, which kept quoting an action item a transcript edit had removed, uses the same counter.
  • Cached models no longer contact the model host when they load — with the text, reranker and visual-model weights already cached, one start still made 25, 32 and 6 requests to the model host for them, carrying a user agent that named the library versions, and retried with backoff when offline; the privacy audit couldn't see it because it sets offline mode for itself. Cached models now load from their local snapshot, and the host is contacted only when nothing is cached (or, for the text and reranker models, when a cached copy fails to load), so the first run still downloads. A cached model never picks up a newer upstream revision on its own. The optional named-entity model, and the visual model with one alternative set of weights, still load by name from the host.
  • The support bundle no longer includes whole folder paths — the path scrubber ran over log lines after JSON encoding, which doubles backslashes and turns non-ASCII into \uXXXX, and it stopped at spaces, so Windows paths, non-ASCII paths and paths with spaces (iCloud folders) went out whole or half-scrubbed. Each record is now decoded, every string scrubbed and the record re-encoded; your home folder becomes ~ first, and paths reduce to <redacted>/<basename> as documented. Prose on a line that also holds a path may lose a few words to the scrubber.
  • Three gaps in the localhost boundary are closed — the Host name the app admits for its own tests (testserver) was also admitted in production, so anyone who could answer your machine's DNS for that single-label name could reach the app from a web page as if it were the app itself, read its data and act with your session; it is now admitted only under MEDIAFIND_TEST_ALLOW_TESTSERVER=1. Cross-origin requests from any page on localhost:8000 were allowed by default; MEDIAFIND_CAPTURE_ALLOW_ORIGINS now defaults to none, and you set it to allow a page. And a request carrying a non-ASCII bearer (with LAN access on) or CSRF token got a server error instead of a refusal.
  • Ask keeps a conversation's scope after a restart — the Ask store kept conversations in memory (lost on restart, after 24 h, or past 200 conversations) while their turns were saved, so a follow-up to an older conversation started over across the whole library and could quote a file you had excluded. The conversation is now rebuilt from its saved turns with its last scope, a reloaded thread restores its scope chips, and a reset stays final. "Verify" no longer calls a quote intact after its recording was trimmed: it checks that the quote's time still lies within its segment, and a receipt with no stored times no longer fails. A Memory reset now also drops the live Ask conversations that quoted the library.
  • Saved searches keep the filters you were using — saving from the app dropped the speaker, person and channel filters, and for Pro users the people channel went missing on every refresh. Saved filters are now validated (a bad shape is refused with 422 and nothing is stored), a save and its first refresh succeed or fail as one, and a re-indexed file stays in the saved searches that held it.
  • Importing or syncing a library no longer makes case-duplicate labels, and deleting the copy no longer deletes the original — import matched collections, tags and saved searches by exact name, while every lookup resolves a name case-insensitively to its oldest row, so deleting an imported "work" tag deleted your own "Work". Names now merge by the libraries' own rule, a clashing saved search comes in as "(imported)" and is reused on the next round, and a merge can't make a collection its own ancestor. Case-duplicate tags and collections an earlier import made are left as they are, and deleting or renaming one of those by name still acts on the older row (saved searches are now addressed by id).
  • "Pull notes from another version" keeps notes where they were — notes snapped to the nearest sampled frame (a note at 607 s landed at 600 s on an hour-long file), and a note that failed to save was counted as reattached. Times are now interpolated between matched frames, and failures are listed as failures. Category chips in search now show their counts.
  • Live meetings (Pro): a meeting paused in another MediaFind process is no longer taken over, and Discard no longer deletes a recording still being saved — a spool untouched for a minute was recovered and removed by any MediaFind process on the same data folder, including a meeting paused in a mediafind meetings live session when the app started; each session now holds a lock on its spool for the spool's whole life (on Windows, the open file itself keeps recovery off). And Discard after a cancelled save could delete the recording and transcript while the save was still running; they are now kept, and the app says so.
  • Video tools: long encodes, cancel, and speed — transcode, compress and GIF had a fixed 300 s ffmpeg limit (a 1080p WebM longer than about 1.6 minutes always failed), left partial files behind and ignored Cancel. They now run with a stall timeout, clean up after themselves and stop when cancelled, including on a file whose length ffmpeg can't read. VP9 encodes are about 7× faster.
  • A file still being written is no longer stamped as indexed — watch took the file's fingerprint after transcription, so a file that grew while it was being transcribed was marked done and kept a truncated transcript for good. Also: an unreadable database gave a misleading migration error and mediafind restore refused to run, and a migration that failed only on a lock still swapped the database file under another process.
  • MediaCreate: imported timelines (Pro) and captions — importing an .otio with a compound clip slid every later clip earlier, so audio lost sync; a camera file with a start timecode imported an hour in; and caption burn-in failed on every Windows install, and anywhere the data folder's path contained :, ', \, [, ; or =. Compounds now take their duration (left as a gap, not flattened), source ranges are rebased onto the media's own start, and the caption path is escaped for both of ffmpeg's parsing levels. EDL and FCPXML exports still write source timecode from zero for a file with a start timecode.
  • A long recording whose timestamps start late is sampled all the way through — since 0.1.46, a recording over eight minutes whose stream starts at, say, 3600 s (DVR and broadcast captures) got one frame, because the seek targets counted from zero. Targets are now offset by the stream's start time.
  • AI clients over MCP and /api/v1/api/v1/transcript for an unindexed file answers 404 with reason: not_indexed (the MCP get_transcript tool returns error: not_indexed), the API's busy 503 says server_busy and keeps Retry-After, and MCP ask_library rejects a blank question. A generated client config names the library (MEDIAFIND_DATA), and Connect now keeps the env you added, including an unquoted "MEDIAFIND_MCP_ALLOW_CONTENT": 0, which used to be dropped on the next Connect, turning content sharing back on; it writes exactly command, args and env and drops stale disabled, type and url keys from another transport. --print-config codex prints env as an [mcp_servers.mediafind.env] table, so a config that already has one still parses.
  • Cleanup: the Duplicates tab is no longer empty once quality flags fill its window — the listing leaves quality flags out unless asked for, reports total and next_offset, and the "Duplicate groups" count no longer includes them. The free tier's 20-group review cap now holds on every page and can't be widened through confidence. Two duplicate scans requested at the same instant no longer both start, and a group dismissed while a scan runs stays dismissed.
  • Also fixed — a search scope of about 16,000 or more files no longer hits SQLite's variable limit; /api/media/{id} reports the file's length, not where its speech ends, and no longer probes the file with ffprobe during an index job; model_required in an Ask answer appears only when a model was needed and none could load; the privacy audit's child process runs with the licence service off, so it can't refresh the subscription or rewrite license.json; a phantom mf-text:// root recorded by an earlier version is pruned; and usage statistics can no longer send counts that a Clear removed while a report was being prepared.
0.1.48
Added
  • Ask your conversation memory a question. mediafind memory ask LIBRARY "QUESTION" and POST /api/v1/memory/{library}/ask answer from one memory library: with the on-device model when one is installed, otherwise by quoting the sentences that match best. The result lists the conversation turns it drew on (add --json on the command line to see them). When nothing relevant turns up, the best matches are weak, or the model declines, you get abstained with a reason instead of an answer. Answers are generated on your machine and never sent to a remote model.
  • Memory search can use the date a question was asked. Give memory search that date (mediafind memory search --as-of 2023-05-30, or as_of in /api/v1/memory/…/search and the MCP memory_search tool). Phrases such as "last week", "in March" or "in the last month" then skip conversations from before that period. Later ones still count, because people often mention something after it happened. A phrase that only sets an end ("before May") or names a period after that date sets no limit. Without as_of, dates in a question are searched as ordinary words.
  • Visual analysis can run in a separate worker process too (experimental, off by default). With MEDIAFIND_MODEL_WORKERS=1, the switch that already moves transcription into a supervised worker process, indexing also runs frame analysis, face detection and image encoding there, so a native crash in them no longer takes the app down. A video that keeps crashing visual analysis is indexed without it, an image that keeps crashing it is skipped, and mediafind index --force retries both.
Changed
  • mediafind share now redacts by default, as the app already did. The command-line share bundle included transcript excerpts, notes and speaker names verbatim unless you passed --redact, while the app's Share (POST /api/share) redacts by default. The CLI now redacts unless you pass --no-redact (--redact is still accepted). Because the original clip still carries whatever was redacted from the text, a redacted bundle leaves it out; add --include-unredacted-clip to include it anyway.
  • mediafind protect now needs Pro, like Protect in the app. It was the only Pro command the CLI didn't check.
  • Command-line flags that were quietly ignored now stop the command. mediafind search --clips refuses --speaker, --person, --no-rerank and --explain, which it used to ignore (exporting clips of every speaker, for example), and it now applies the Pro check for face and logo search. ask --across --clips, which exported nothing, and highlights with both --query and --tag are refused too. collections move NAME with neither --parent nor --no-parent used to clear the parent; it now asks for one. mediafind index and mediafind watch --once on a path that doesn't exist exit with an error instead of reporting "Done. files=0", and both now expand a quoted ~. audiogram, share, file, meetings process and several notes and transcript subcommands print failures to stderr and now exit non-zero.
  • Memory search finds the right conversation more often by default. Searches of conversation memory (mediafind memory search, the MCP memory_search tool and /api/v1/memory/…/search) now use the lexical mode unless you pick another. The previous default, rerank, runs a reranker trained on web passages over chat turns. On the LongMemEval-S benchmark, the conversation holding the answer was among its ten results for 79% of questions, against 94% for lexical. --mode rerank is still available.
  • The knowledge map (Pro) stays readable on big, densely linked libraries. When most files shared a speaker, a face or a category, both link modes (shared signals and semantic relatedness) drew every link at once, so the map became a single-colour tangle with file names piled on top of each other. The map now keeps each file's three strongest links and groups files into topic clusters that sit apart from each other, each named, where it can be, after a category that at least half of its categorised files share. It labels the clusters rather than every file, and shows file names as you zoom in. Hovering a file still reports all of its connections, and the legend now sits under the map instead of covering it. Semantic-map nodes from GET /api/knowmap now carry their categories, as shared-signal nodes already did.
Fixed
  • Emptying the quarantine no longer de-indexes a file that is back in place — Empty quarantine, and deleting a single quarantined item, removed the index entry at the item's original path without checking what was there. A file you had dragged back out of the quarantine folder in Finder instead of clicking Restore, or a different file saved at that path later, lost its transcript, notes, tags and collection memberships while staying on disk, and its size was reported as freed. A file that is back at its original path now stays indexed.
  • Re-transcribing a file no longer deletes its meeting notes — a forced re-index (Re-transcribe, including re-transcribing everything after you switch the speech model) deleted the file's processed meeting: its overview, decisions and action items, ticked-off ones included. They now stay with the file.
  • Running the privacy audit no longer breaks search and indexing — while the audit ran from the app (GET /api/audit), every other search and indexing job in the server used a stand-in embedder, so files indexed meanwhile were stored with stand-in vectors and searches came back empty. Two audits at once could leave the server in that state until it restarted, and an audit early in a session could make later model downloads fail as "offline". The server now runs the audit in a separate process, one at a time, and refuses to start it for another website open in your browser. mediafind audit is unchanged.
  • A live meeting that fails to start no longer leaves the microphone on (Pro) — if anything failed after the browser granted the microphone or tab audio (for example Firefox refusing to connect a 48 kHz microphone), the capture kept recording with no way to stop it, and every further Start opened another one. The app now stops the capture and says what went wrong, and a microphone that refuses 16 kHz is retried once at the browser's own rate. Discard also stops at once now: it used to transcribe the audio still queued, and add those voices to your voice library, before discarding.
  • Names you give people and voices no longer move to someone else (Pro) — labels such as "Person 3" and "SPEAKER_00" can be freed and given to someone new, and several paths left your names and confirmed voice links pointing at a label after it went to a different person: re-grouping faces after a person's media was removed, clearing out people and voices with no media left, and diarizing a file whose voices matched no one, which could file a new voice under whoever the library already called SPEAKER_00. Merging two speakers now keeps the person link, a voice you linked survives the removal of its recordings, and names are stored without stray spaces, so "Bob " can be found as "Bob".
  • Speaker diarization tells voices apart again in new installs and the Windows app — the voice model behind diarization needs a voice-detection library, and that library still imported pkg_resources, which setuptools removed in version 82. An install or build that picked up a newer setuptools (torch pulls it in) could not load the voice model, so diarization fell back to MFCC features that can't separate similar voices, and a two-person recording could come out as one speaker. That hit fresh pip installs and the Windows build, and a Mac app built the same way did it too. MediaFind now loads that library without pkg_resources, and the app builds stop if the diarizer can't import its voice model. Recordings diarized while the fallback was in use keep the speaker labels it gave them, and MediaFind can't recognise their voices in new recordings, until you run speaker detection on them again (Detect speakers, or mediafind index <folder> --diarize --force).
  • mediafind doctor notices when speaker diarization can't import its voice model — it only checked that the model's package was installed, so an install where it failed to import still reported the diarizer as ok, even in the release builds' own checks. doctor and those checks now import it for real. And once a running app has fallen back to the older MFCC voice features, its diagnostics and live meetings report that instead of promising speaker recognition.
  • Importing a conversation no longer replaces a different one with the same name — importing 2024-02/team-sync.txt into a memory library that already held 2024-01/team-sync.txt replaced January's conversation, and the CLI reported both as imported. For conversations imported from this release on, a different file with the same name is now refused, naming the file already there; re-importing the same file, edited, still updates it. Resetting a memory library now also deletes the Ask answers that cited it.
  • Healthy files are no longer quarantined in place of the one that crashed — while several files were being transcribed at once (the default with 8 or more CPU cores), a crash in one was blamed on another, which after three crashes was skipped as a "poison" file. Quitting the app or pressing Ctrl-C while a file was indexing also counted as a crash, so three quits could get that file skipped. Files that were in flight during a crash are now retried one at a time, so a healthy neighbour isn't blamed again and skipped. A long file transcribed alongside others is also no longer failed by the two-hour no-progress watchdog, and a free-tier batch that hit the file limit and also had a failed file now shows the limit.
  • Usage statistics count what you actually did — the local counts behind the opt-in usage reports were off: paywall hits, active days and the first-paywall milestone were counted from the app's own background refreshes, and a mediafind watch --once run from cron marked the day as active. Demo mode was counted, and those counts could go out with the first report after demo mode was turned off. Clearing the statistics could let a day's once-only counts be sent twice, a Clear made while a report was being prepared could let the cleared counts go out, and imported conversations counted as media files in the library size. Demo mode now counts nothing, and the rest are fixed.
  • A wrong clock at one launch no longer ends the free trial — a clock reading in the past (a dead clock battery, starting up before the time synced, a restored virtual machine) looked like a tampered trial start, and the defence against that moved the trial's start for good, so the trial counted as expired once the clock was right again. MediaFind now remembers the latest time it has seen and tells the two apart, from the first launch of this version with the clock right.
  • Libraries exported without embeddings can be searched after import — such an export still recorded which model had made the embeddings it left out, so the imported library looked embedded: search by meaning found nothing in it, and refreshing stale embeddings skipped it. New exports without embeddings are marked as not embedded, so refreshing embeddings after import covers them. Export also checks the importer's size limits now (raised from 50 MB to 512 MB for the library's JSON files) and fails up front, instead of writing an archive that fails only when you restore it.
  • Search no longer fails on big folders or long queries — a search limited to a folder or a filter covering 500 or more files returned a server error, and a query of a thousand or more different words, such as a pasted document, could fail the same way.
  • Visual citations no longer all say "source changed" — every Ask answer that cited a video frame or on-screen text showed "source changed" for good, because such a citation has no transcript line to check. They are now checked against the cited file's recorded size and modification time, so they flag only a file that changed since the answer.
  • Live meetings (Pro): a pinned language holds, a failed save can be retried, and a crash no longer strands the audio — with the Apple GPU transcription engine, a meeting pinned to a language (Chinese, say) was transcribed as English, or in the language of the last library file. A save that failed (on a full disk, for example) could never be retried, leaving the recording without its meeting. After a crash mid-meeting, the raw audio sat in a file nothing read; the app now saves it as recordings/meeting-<time>-recovered.wav when it next starts at least a minute after the crash.
  • More live-meeting fixes (Pro) — Stop & save no longer reports the meeting as "stopped elsewhere". Exported notes and saved recordings keep non-Latin titles ("产品评审会") instead of a code or a mangled name, and a speaker name containing a backslash no longer breaks the live summary. mediafind meetings live now keeps recognised voices and the AI summary, as the app does, and with MEDIAFIND_CAPTURE_SR set to 48000 the server microphone no longer records audio that plays three times too slowly.
  • AI summaries work for long Korean meetings and recordings — MediaFind underestimated how many tokens Korean, Japanese kana and Cyrillic text takes, so from about 14 minutes of Korean speech every live summary pass failed, and long Korean files got no summary. English is unaffected.
  • Ask keeps working while another model downloads — while the model picker downloaded a tier, Ask on the model already loaded, live-meeting summaries, summaries made during indexing, and skipping the download in onboarding all waited for the whole download to finish.
  • Action-item due dates land on the right day — in the follow-ups list and the calendar (.ics) export, "by Friday next week" meant this Friday, "by the 15th of next month" this month's 15th, and "by Monday, October 5th" the coming Monday, while the end of "the week after next" or "the month after next" became this week's or this month's. These now resolve correctly, or stay undated when the phrase can't be pinned to a day. Task owners are no longer word fragments ("Land" from "notify the Landowner"). Meetings processed from now on get these right; process an existing meeting again to redo its tasks. mediafind meetings ics now dates tasks from the day the meeting was processed, as the app does (--ref still overrides).
  • Duplicate review does what it says — for two identical copies, "Kept the earliest copy as the likely original" could keep the newer one; the oldest copy now wins when nothing else tells them apart, and the reason says "earliest" only when it is. For groups dismissed from this release on, "Never suggest again" also holds when the group shrinks, for example after you delete one of three copies. Starting a duplicate scan while one is already running is refused, because two at once could drop real pairs from the review queue; a group dismissed during a scan stays dismissed; and a cancel at the very end of a scan no longer leaves the results half replaced. On very large libraries, the similar-image and similar-video scans (Pro) stop at a work budget and return partial results instead of running out of memory.
  • Maintenance jobs no longer offer a Cancel that can't work — rebuilding the vector index, refreshing embeddings, cleaning up thumbnails and re-grouping people can't stop partway, but offered Cancel anyway: it marked the job cancelled while the work carried on, and let a second copy start. They no longer offer it. Two requests to rebuild the vector index or refresh embeddings sent at the same moment can no longer both start, and quitting with jobs queued leaves them queued for the next start instead of running the whole backlog first.
  • Transcripts made by a fallback model can be re-transcribed — when MediaFind stood in another model for the one you picked (a bundled model while yours downloads, or the default engine for the Apple GPU one), it recorded the one you picked, so the stale-transcript check never offered to redo those files. It now records the model and engine that ran. Diarization likewise records when it used the MFCC fallback, so a later re-index with speaker detection redoes those files once the voice model is available; files diarized before this release are left as they are.
  • Imported conversations stay out of places meant for media — they no longer show up as file or folder matches in search, in the search Folder list, in the home "On this day" rail, or in re-transcribe all. Exporting a conversation as text or Markdown from the app now works (it said the file was no longer on disk), and subtitle formats are refused, since turns have no media times. Plain transcripts accept a speaker name alone on its line, a file saved with a byte-order mark keeps its first speaker and time, and times ending in "Z" are read correctly on Python 3.10.
  • Memory searches limited to a date range find more of the turns inside it — the after/before range was applied to the top results, so when stronger matches fell outside it, a search could return fewer turns than asked for even though more inside the range matched. The range is now applied during the search.
  • Model workers (experimental, MEDIAFIND_MODEL_WORKERS=1) — a short file queued behind long ones failed as a timeout, because its deadline started before a worker picked it up; a model build that ran out of time could leave indexing spinning; cancelling while a video was transcribed in the worker saved it with an empty transcript; and once repeated crashes had turned transcription off, switching the model or engine didn't turn it back on, and videos indexed without a transcript reported plain success. All fixed; the workers stay off by default.
  • Chat-control markers in transcripts, notes and conversations no longer reach the on-device model's prompt. Markers such as <|im_start|> in transcripts, notes or imported conversations are now neutralized before they reach the local model. This applies to Ask as well as to memory answers.
  • No more segfault when the text model loads after the visual model with pyarrow installed. On macOS, in a Python environment that also had pyarrow (the benchmark extras install it), loading MediaFind's text model after its visual model killed the whole server or CLI process with a segmentation fault. pyarrow's bundled memory allocator crashes when a thread first allocates after the thread that loaded pyarrow has finished, and MediaFind loads each model on its own short-lived thread. MediaFind now has pyarrow use the system allocator (ARROW_DEFAULT_MEMORY_POOL=system) unless you set that variable yourself; a program that imports pyarrow before MediaFind should set it too. The desktop app doesn't include pyarrow and was not affected.
0.1.47
Added
  • Indexing now steps aside while you work, on macOS. Background indexing pauses between files while a pro or creative app is in front (Final Cut Pro, Premiere Pro, After Effects, DaVinci Resolve, CapCut, Logic Pro, Ableton Live, Pro Tools, Blender, OBS Studio and others), while your Mac is running hot, while memory is running low, or, when MediaFind uses the GPU, while another app keeps it saturated for a minute. It carries on by itself when that clears. The activity dock says why ("Paused — Final Cut Pro is in front"), and a paused job can still be cancelled. When memory is only under pressure, indexing keeps going at a lower priority instead of stopping. This is on by default: MEDIAFIND_YIELD=0 (or the yield_to_apps preference set to false) turns it off, and MEDIAFIND_YIELD_APPS adds apps by bundle id.
  • Conversations and notes as searchable memory (preview). MediaFind can now take in chat transcripts and notes alongside your recordings. On the command line (Python installs), mediafind memory ingest reads a ChatGPT or Claude data export, MediaFind's own JSON, or a plain speaker: text transcript, and mediafind memory search returns the matching turns, each with who said it and, when known, when. An AI assistant can keep its memory in MediaFind through three new MCP tools (memory_save, memory_search, memory_reset) or the new /api/v1/memory routes. What it stores is a library you can list, search and reset yourself (mediafind memory sessions, search, reset). Conversations are stored and embedded on your machine; a connected agent sees the text it recalls, and MEDIAFIND_MCP_ALLOW_CONTENT=0 withholds memory text from agents too, and mediafind audit now covers the conversation path. Conversations don't count against the free tier's file limit.
  • Opt-in anonymous usage statistics; sharing is off by default. MediaFind now keeps daily counts of what you do in a local usage.sqlite3: searches run and whether they found anything, results opened, exports, and features used. If you turn on *Settings → Usage statistics* (or run mediafind usage --share on), the app sends those counts every few hours so we can see which features help, with the app version, platform and install week, a weekly note of your plan and your library's size and media mix, and how long after installing you reached a few milestones, in coarse buckets. Reports never contain file names, paths, search text, transcripts, faces, or any persistent user or device identifier, and you can preview what a report contains before anything is sent. mediafind audit still shows the core path opening zero external sockets.
  • Transcription can run in a separate worker process (experimental, off by default). With MEDIAFIND_MODEL_WORKERS=1, indexing transcribes in a supervised worker process, so a native crash in either transcription engine no longer takes the app down with it.
Changed
  • Searching within a folder, a file or one speaker's lines is faster on large libraries. A filtered search used to build a full record for every matching passage before keeping the best few. It now scores the passages first and reads only the ones it returns. On a 50,000-passage scope, building those records was 116 of the 165 ms a search took. Scoped Ask questions and the speaker filter get the same speed-up.
  • The logo and action lists open without loading the visual model, and indexing no longer loads models just to name them — the first visit to either list after launch loaded the visual model only to report whether it was available: 8.3 s with the model cached and 14.4 s without, in a fresh process. It now takes about 0.01 s. Indexing no longer loads the visual, text, face or entity models just to name them in its summary, so a re-index with nothing to do no longer pays for loading them, and looking up names in the celebrity gallery no longer starts face recognition.
  • Scene searches are faster on large video libraries — a query for a scene the model knows by name, such as "kitchen" or "office", now scores every frame at once instead of one frame at a time. Scores and ranking are unchanged.
  • The app reaches its window sooner — launching no longer loads PyTorch or OpenCV before the server can answer, since many sessions never need them; they load when a search or an index job first does. MEDIAFIND_STARTUP_TRACE=1 logs the timing of each startup phase.
  • Searches from AI agents and the command line leave room for other work — the MCP server and mediafind search / mediafind ask now cap PyTorch at half the CPU cores by default, as the app already did (MEDIAFIND_TORCH_THREADS overrides it).
  • Release — 0.1.46 was built but never published as a download, so this release also carries everything listed under 0.1.46 and 0.1.45. It ships a signed, notarized macOS DMG and native Windows and Linux beta installers.
Fixed
  • The Windows and Linux downloads now include visual search, face recognition and fast vector search — their installers and server binaries were built without the visual model, the face model and the ANN index, so visual, logo and action search and People & Faces did not work there. They now ship them, with the face-recognition weights bundled as on the Mac, along with the media downloader, HEIC images, editor-timeline export and offline license-key checks. People & Faces, logo detection, the downloader and editor-timeline export are Pro features.
  • MEDIAFIND_MCP_ALLOW_CONTENT=0 now keeps library text out of MCP search results too — the switch made ask_library and get_transcript refuse, but search_media still returned each hit's transcript text, summary, on-screen (OCR) text, matched note or chapter title, and names spoken in it, so a connected agent could read the library through search anyway. With the switch off, search still finds moments but returns only where each one is and how well it matched: file path, timestamps and scores. Labels such as speaker and person names are withheld too, because imported transcripts and detection files can put any text in them. With the switch on (the default), search is unchanged.
  • A file's Apple GPU engine transcript no longer depends on what was transcribed before it — re-indexing with the Apple GPU engine could give a different transcript each time. When the model re-decodes a stretch by sampling (music, near-silence, and about one ordinary file in five in a test set), it draws from a random stream that the engine seeds once and never resets. So one clip came back as 14 segments when it was transcribed first, and as 19 with different words after a music video. With several indexing workers, the order (and so the text) changed from run to run. Each decode now starts from a freshly seeded stream: after a decode that sampled, the model is reloaded from its weights file, which takes about half a second (up to 3 s on a busy machine).
  • A file's default engine transcript no longer depends on what was transcribed before it — when the model re-decodes a stretch by sampling (music, near-silence, some speech), the engine drew from a random stream that each of the model's worker threads seeds once, from the operating system, and never resets. So a file's transcript depended on what the app had transcribed before it, on which indexing worker took it, and on the run: one music-vlog clip came back as three different transcripts from three decodes in the same run. Each stretch that samples now does so on a freshly loaded copy of the model with a fixed seed, which holds a second copy of the model in memory while the stretch decodes; if that copy can't be loaded, the stretch samples on the shared model as before. Turning the sampling off instead was measured worse: stretches of music became one phrase repeated dozens of times. Transcription on an NVIDIA GPU is unchanged.
  • Meeting notes missed commitments written with a curly apostrophe — a transcript that spells "I’ll", "let’s" or "we’ve" with a typographic apostrophe (imported subtitles, sidecar transcripts, pasted text) matched no action or decision cue. So "I’ll send the deck to legal by Friday." was never listed, and "I’ll set up the call with Priya, please…" went to Priya instead of the speaker. Cue matching now treats ’, ‘ and ʼ as a plain apostrophe, and the stored task text keeps the transcript's own spelling.
  • An import could deadlock inside numpy's math library, and searches and the live panel waited on its thread pool while indexing ran — numpy's bundled math library (OpenBLAS) kept a pool of every core. When one thread started a helper program (indexing runs ffprobe for each audio and video file) while another was inside a multithreaded calculation, which indexing's look-ahead workers and a concurrent search both do, the process deadlocked: indexing stopped making progress forever, with no error, while keeping every core busy. Under load, each small calculation on a search or in a live meeting's panel also waited for a worker per core. OpenBLAS now runs on one thread; an explicit OPENBLAS_NUM_THREADS still wins. Measured under load, scoring a search against 100,000 passages went from 76 ms to 5 ms; clustering the largest face libraries (20,000 faces) takes up to about 10 seconds longer.
  • Due dates like "March 3rd" and "the end of the week" — in a meeting's calendar (.ics) export and in Follow-ups, "on the 3rd of March" was dated the 3rd of next month, and "by March 3rd", "15 June" or "by the end of the week" got no date at all. A month date now resolves to the next such day on or after the meeting, and "by the end of the day / week / month / quarter / year" to the last day of that period (a week ends on Friday; quarters are calendar quarters). A close-of-day cue keeps the day it names: "by end of day Friday" or "by EOD next Friday" is due that Friday. A date with a year is that year's ("by March 3rd, 2028", "by October 3rd next year"); one already past, or a cue that can't be read, stays undated rather than guessed.
  • Music and near-silent files no longer keep a transcript stuffed with one repeated line — when the voice gate leaves a file almost empty (it scores sung vocals as non-speech), MediaFind transcribes it again without the gate. That second pass can lock onto a phrase and repeat it dozens of times, often after real speech, and a transcript that was about half loop was kept whole, loop and all. Segments made only of a phrase of up to 16 words said four or more times in a row are now cut out, and what remains is kept only if it still passes the existing checks on its own. A second pass left with fewer than 20 words, such as "Thank you." five times over silence, is no longer kept at all.
  • The 100 MB request-size cap now also covers bodies sent in chunks — the local server checked a request body against the cap (MEDIAFIND_MAX_UPLOAD_MB, 100 MB by default) only through its Content-Length header. A body sent in chunks carries no such header, or a small one the server then ignores, so every route that takes a JSON body (most of the API) read the whole body into memory, however large. The server now counts the bytes as a route reads them and answers 413 as soon as the total passes the cap, with the same error an oversized Content-Length already got.
  • Cancelling a job just as it started could say it had failed — if the job was starting, or resuming from a pause, at the moment you cancelled it, the app said "This job cannot be cancelled." although the job stopped. That cancel now reports success, and a cancelled job can no longer be marked as failed if the job worker hits an error afterwards.
  • Finder showed version 0.0.0 for MediaFind.app — Get Info and crash reports gave every release so far as 0.0.0. The app now carries its real version.
0.1.46
Added
  • Live meeting (streaming) — transcribe a meeting *as it happens*: audio streams to the local server and a rolling transcript with provisional, renamable speaker labels plus live notes (key points, action items, decisions) updates in real time. Stop saves the recording as a WAV in the library and files it as a normal Meeting. Sources: browser mic, tab/screen audio for online meetings, or the server's own mic via the [meetings] extra. Python installs also get it on the command line as mediafind meetings live. Keyless and on-device; Pro.
  • Live meetings recognise voices you have named — speaker labels in a live meeting are matched against your existing on-device voiceprint library, so someone you have already named is shown by name from their first sentence. Renaming a speaker during a meeting remembers that voice, and the next meeting identifies them automatically. Matching only runs on the real voice embedder; on the signal-only fallback a meeting keeps session-local "Speaker N" labels rather than filling the library with voiceprints that could never match. Local and keyless, like the rest of Meetings.
  • Live AI summary and visual summary — with a local language model installed, a live meeting shows an AI-written summary that streams in as the meeting goes, next to a visual summary: a topic timeline, each speaker's share of talk time, and the key terms, ranked by how distinctive they are against your library rather than by raw frequency once it has indexed files to compare with (character-pair terms for Chinese, omitted when a short meeting has none that repeat). When search has already loaded the text embedder, the topic timeline uses it, so two subjects that share a word still come out as two bands; a meeting never loads a model itself. The AI summary is model-written and can get details wrong — a small local model can swap a date or an owner — so it is labelled as such and kept apart from the extracted action items and decisions, which are read straight from the transcript and stay the source of truth. The summary covers the whole meeting: once a meeting outgrows what the model can read at once, each pass updates the previous summary with what was said since. It is built on those extracted action items and decisions, and the model is told to keep their owners and due dates exactly as the notes give them. A sentence stating a date or number that nobody said is dropped (dates are read in English, plus Chinese weekdays; other languages are checked for digits only). With no model tier chosen in Settings and the model on the GPU, live summaries use the largest model already downloaded, up to Medium. Streaming holds back the last few words until the model's text is vetted, so a model reciting its own instructions is caught before any of it is shown. Without a local model the panel keeps the keyless notes and says a model is needed; nothing leaves the machine either way.
  • Saved live meetings keep what you saw — Stop now files the meeting with the AI summary you watched as its overview, instead of generating a second, different one. One closing pass while it saves brings that summary up to the meeting's last line: a live pass runs every 45 seconds, so the summary on screen usually misses the last minute. Each voice you named is stored under its voice-library entry, so the saved meeting joins People & Voices like any diarized recording and a later rename reaches it. Previously every named voice also minted a voiceless duplicate speaker.
  • Live meetings in the desktop app — the Microphone source now works inside the app window: WebKit's own permission sheet used to leave it waiting forever. macOS still asks once whether MediaFind may use the microphone, and signed builds now carry the audio-input entitlement the hardened runtime requires.
Fixed
  • Batch de-dup could leave you with no copy of a file — Cleanup's Apply could move every copy in a duplicate group to quarantine when the copy it meant to keep had been removed from the library or sat on an unplugged drive, or when that copy was also being removed as part of another group. Groups whose keeper is missing are now skipped and listed for review, and a file kept in one group is never quarantined as part of another.
  • Restoring the oldest backup could empty the library index — the safety backup taken before a restore pruned the very backup being restored, and the restore then copied an empty database over the live index while reporting success. The backup being restored is never pruned now, and a missing backup fails the restore instead of being replaced by an empty file.
  • A file restored from quarantine could join an unrelated duplicate group — marked as a high-confidence exact duplicate, where the next batch clean-up could quarantine it again.
  • Re-indexing a file wiped its notes, tags and collections — a re-index deleted and re-created the file's record, and the notes, tags, collection memberships and a download's source link went with it while the job reported success. They now carry over. Four more ways of losing data are fixed too: a re-index that fails part-way no longer removes the file's record, a read error no longer lets thumbnail cleanup delete everything it references, a failed or cancelled duplicate scan keeps the previous results, and one malformed transcript segment no longer loses the whole file.
  • Long videos were only sampled for their first 8 minutes — frame extraction stopped after 240 frames at 2 s apart, so anything later in a video was invisible to visual, face, on-screen-text and every other frame search. The same frame budget is now spread over the whole video. Re-index a long video to add the rest.
  • Redacted exports still carried the spoken words — the transcript file kept the text of every word in its per-word timings, and the summary, chapter titles and on-screen text were left as they were. Redaction now replaces all of that text while keeping timings and speaker labels, and a shared redacted bundle no longer sits in a folder named after the original file.
  • Clip exports now refuse requests from other websites.
  • Re-running a backfill could empty its channel — the audio-event, object, colour and entity backfills cleared the stored results first, so quitting or a crash while they recomputed left the channel empty. They now compute first and only then replace the old results.
  • Regrouping faces could put a person's name on someone else — when faces below the minimum size were dropped during regrouping, a named person left with no faces handed their name to the nearest other group.
  • A restart could hide stored video frames from index rebuilds — rebuilding the visual search index, or re-running detection, right after launch could skip every stored frame.
  • Folder-scoped search missed matching results on the action, logo, scene, audio, emotion, phonetic and object channels. Results were cut to the library-wide top matches before the folder filter ran, so a scoped search could return nothing while matches existed.
  • Some status messages reported work that never happened — Ask's evidence note said visual coverage was complete when no visual evidence existed, the assistant showed "Done" for a highlight reel that rendered nothing, event counts claimed full coverage for detectors that had never run, and a task could be logged as executed when it changed nothing.
  • Cleanup kept offering space it had already reclaimed — files moved to quarantine were still counted as reclaimable duplicates.
  • Connect to AI agents now works on the Windows and Linux builds and from source installs — those builds were missing the MCP server package, and for pip or source installs the configuration written for your AI client named mediafind-mcp without a path, which desktop AI clients could not find. The signed macOS app was not affected.
  • Also fixed — search could stop responding until a restart after one failed database reopen; paused indexing jobs could leave a worker stuck and delay quitting; switching the model device in Settings could crash a search in progress; Home's People count stopped at 500; the Ask backend chosen in Settings was ignored; follow-up questions in Ask skipped the safeguard against instructions hidden in your media; exact-name search matched inside longer names (a search for "Ali" found "Natalie"); Meetings could stop refreshing; the Apple GPU engine could stamp a segment across a silence it had cut out; a very large start time on a title or lower-third could break a project's GIF export for good; two processes rebuilding the same library's vector index at once (for example the app and the CLI) could collide; and a verbatim quote that semantic search scored low could lose its place in the results.
  • The Large and Ultra AI-assistant tiers could never download — both pointed at a single model file that does not exist upstream (those two models are published split into 2 and 3 shards), so choosing either tier failed with a 404. They now point at single-file models of a newer generation (see *Changed*).
  • A downloaded model tier could not load without a network — the AI engine's loader lists the repo on the model host before it looks at the cache, so with no network every tier except the bundled Mini failed to load, even on the offline (local_files_only) summary path. Weights now resolve through the local model cache first and download only when missing.
  • Chapter titles, task routing and caption translation got a raw completion — every tier is an *Instruct* model, but only summaries went through its chat template; everything else made the model continue the prompt as a document. Measured on 39 Ask questions plus 17 probes of the other callers: LLM chapter titles came back empty on the smaller models (0 of 4 → 4 of 4), the task router picked the right recipe 1 of 6 times on Small (→ 5 of 6), and answers dragged the cited context along (avg 154 → 59 characters on Small). Every call now uses the chat template.
  • Ask could hand back its own safety instructions — the bundled Mini model sometimes repeated the instruction block it was given instead of answering, cut off mid-word. Ask now rejects any answer that repeats those instructions and falls back to the cited extractive answer.
  • The search trace loaded a model, and described a pipeline the caller could not runGET /api/search/trace exists purely to narrate how a query is retrieved, and both its own docstring and search_trace.py's promised that no model is loaded. It probed the visual-model backend with visual.is_semantic(), which resolves through get_encoder() and *builds* the encoder, so asking for an explanation could start a multi-gigabyte weight download — a hang, not a delay, on a slow network. The backend is now observed without constructing it (visual.loaded_backend(), falling back to the capability probe when nothing is loaded yet, so a cold process is neither over- nor under-claimed). The same endpoint also skipped the entitlement projection /api/search applies: a free user asking for channels=people got a 200 trace describing a people retrieval step their search would have refused. The trace now runs the same projection and returns the same 402. Finally, an unqualified "all" query is resolved through search._norm_channels like the real dispatch, so the trace stops naming channels (people) that an "all" search never runs.
  • File summaries showed the prompt instead of the file — every model tier is an *Instruct* model, but summaries were generated with a raw completion, so the model continued the prompt rather than answering it. On a 200-file library 57 of 191 stored summaries (29.8%) contained prompt text (32 the literal <context>), the median summary was 1,321 characters against a prompt asking for "2-3 sentences", and 158 of 191 (83%) ended mid-sentence at the token cap. Summaries now go through the model's own chat template, are trimmed to the requested number of complete sentences, and drop any sentence that quotes the prompt back. Re-measured on 15 of those files with the bundled Mini tier: 0 of 15 leaked prompt text (was 4 of 15), median length 293 characters (was 1,350), 1 of 15 ended mid-sentence (was 12). Existing libraries are re-summarised by Backfill summaries, which needs no re-indexing.
  • No crash on exit after using the local model — with the text embedder on the GPU (the default), a process that had run the local model aborted at shutdown (exit 134): a correct CLI answer followed by a native crash, and a crash report when the desktop app quit. The model is now released cleanly before exit.
  • Meeting briefs show speaker names — action-item assignees and decision speakers showed raw labels such as SPEAKER_03 for voices you had named; the brief now resolves them the way the follow-ups list always has.
  • Meeting notes: due dates, false action items and owners — an action item's due date now always names a time: "can you run a usability test on the draft before Monday?" was due "on the draft", and "by the way" or "on 3 accounts" counted as deadlines. Lines that only run the meeting, such as "let's start the review" or "let's go through the checklist", are no longer listed as action items, while "let's send the deck to legal by Friday" still is. After "Yes," or "Okay," an "I'll…" commitment now belongs to the speaker instead of the next capitalized word ("Friday", "Priya"). Applies to the live notes and to every processed meeting's brief.
  • Cancelling a job could say it had failed — when a running job noticed the cancel request before the request finished, the app said "This job cannot be cancelled." although the job did stop. That cancel now reports success.
Changed
  • Local AI-assistant tiers move to a newer generation of on-device AI models: Mini (bundled), Small, Medium, Large and Ultra all get new models. The old Medium's model is under a non-commercial licence. The hybrid-reasoning tiers render their chat template with thinking off, so answers start on the answer, not a <think> block. Ask's prompt now asks for "I don't know" — the reply its refusal check recognises — rather than "say you don't know", which the new models took literally ("You don't know. …"). On 7 questions the library cannot answer, the tiers stopped inventing an answer far more often: Mini 0 → 2 honest, Small 1 → 4, Medium 5 → 7, and Large and Ultra 7 of 7 (both previously undownloadable). Someone already on Small keeps using their downloaded older model until the new model is downloaded from Settings; Medium re-downloads. The llm extra now requires the first release of the on-device AI engine that can load the new models.
  • Large libraries search faster — the check that decides whether the fast vector index is still fresh is far cheaper and remembers its answer, folder filters are compiled once per search, on-screen-text and filename lookups are much narrower, and the database health check is cached. On a copy of a real library scaled 20×, a warm multimodal search went from 2.4–2.8 s to 1.1–1.8 s, and rebuilding the vector index no longer blocks searches while it runs.
  • Less background work — the Knowledge map no longer redraws continuously on every page, a visual-verification check could poll forever, and indexing no longer refreshes the whole app after every file; while indexing, Home refreshes about every five seconds instead.
  • Visual and text search run on the Apple GPU by default — measured on an M-series Mac, visual-model image and text encoding and the transcript embedder are 3.5–4.3× faster. CPU and GPU embeddings agree (cosine 1.000000), so nothing needs re-indexing. Settings → "Visual & text model device" switches back to the CPU.
  • Searches that don't pick channels now include actions and scenes — the default for the CLI, the MCP server and API calls without an explicit channel list.
  • Sharper ranking — transcript hits point at the matching moment rather than a whole 10–40 s transcript block, names the app recognises are left out of what the visual model is asked to match so it can focus on the rest of the description, quotes that span two transcript segments match as one, and a visual hit ranks higher when the neighbouring frames match too.
  • Transcription keeps more sung vocals — music that the speech filter partly rejected is retried, so live performances keep more of their lyrics. Existing files need re-transcribing to benefit.
  • The Apple GPU engine now matches the default engine — it filters silence the same way, produces word timestamps and sets segment boundaries from the words. Face recognition and the on-device language model still default to the CPU; environment variables can move them to the GPU.
  • Release — 0.1.45 was built but never published as a download, so this release also carries everything listed under 0.1.45. It ships a signed, notarized macOS DMG and native Windows and Linux beta installers.
0.1.45
Fixed
  • The People grid fronted a person's worst crop — the representative face was the detection with the highest *detector confidence*, which barely tracks size and ignores blur. On a real library half of all people were fronted by a crop under 48px, e.g. a 22px smear chosen over a 122px face of the same person. Every face now stores a crop-quality score (size × sharpness × confidence) that picks the representative in the People grid, file chips, face strip and appearance segments. Existing libraries are backfilled from box size on first open; Re-detect faces recomputes with sharpness.
  • Postage-stamp faces polluted people — detections under 24px on the short side (12% of that library; 13% of its "people" had no larger face) are blobs with no identity signal, so the live detector now skips them (MEDIAFIND_FACE_MIN_SIZE, 0 to disable). Face crops are stored at 192px instead of 160 so a 72px tile on a Retina display no longer upscales them.
  • One person kept splitting into several — face-level average linkage at the 0.55 threshold separates a person's pose and lighting variants: on a real library many of the names the keyless gallery matched were split across several "people" (Elon Musk across 8, Donald Trump across 11), with the fragments' centroids only 0.44 apart while different people's centroids never came within 0.81. Clustering now runs a second, centroid-level pass (MEDIAFIND_FACE_MERGE_DIST, default 0.45, 0 disables) that folds such fragments together: replayed on that library it put the fragments back with their people, cut the split names by more than half, and merged no two different names.
  • Regroup kept re-creating junk people from already-stored tiny faces — a library indexed before the size floor still holds its sub-24px detections. Re-clustering (Regroup, mediafind people --cluster, and the incremental pass after every index run) now leaves those faces ungrouped instead of spawning a one-face "person" each, and the file view no longer nags to refresh faces over them. Re-detect faces drops them for good; MEDIAFIND_FACE_MIN_SIZE=0 groups them anyway.
  • Place words no longer hijack ordinary search into a same-named collection. A query such as "sunset in Tokyo" could be hard-filtered to a collection named Tokyo even when that collection did not exist, hiding otherwise valid results.
  • One-frame people appearances survive zero-padding. Appearance segments that begin and end on the same frame are now retained instead of disappearing when the configured segment padding is zero.
  • Production builds ignore local licensing overrides. Frozen apps no longer honor the development-only environment switches that unlock Pro; the free and trial paths now run in CI so this boundary cannot silently regress.
  • Local Ask no longer repeats its hidden instructions after the answer. Echoed prompt text from the bundled on-device model backend is removed before display.
  • Knowledge-map links render in both supported modes. Switching between the entity graph's link views no longer leaves the map with missing connections.
  • Recovery and accessibility paths are reliable again. Restores no longer race the database pool, paused indexing resumes correctly, keyboard and screen- reader flows pass their UI checks, and the compatibility matrix reports the tested platform support accurately.
Changed
  • People page at scale — a 16k-person library shipped its whole people list (4 MB) in every full status load and rendered every tile at once (108k DOM nodes, a 16k-option person filter). The status payload now embeds the top 500 people (MEDIAFIND_STATUS_PEOPLE_MAX) plus the true total, the grid reveals 120 tiles at a time behind a Show more button that fetches the complete list only when needed, and the person filter offers the 300 most-seen people plus the current selection.
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current main release head with synced 0.1.45 package, desktop, MCP bundle and storefront metadata.
0.1.43
Fixed
  • ASR: a file with no audio track is no longer reported as a decode failure. The default engine indexes container.streams.audio[0] unguarded, so a video-only file raised a bare IndexError that surfaced as "decode failed" and left the file queued for a retry that could never differ. Now a typed NoAudioStream, converted on both decode paths.
  • ASR: sung vocals are no longer silently dropped. Voice detection scores singing as non-speech, so music videos were indexed with an empty transcript and no trace of why — on one file the gate removed all 3m20s of a 3m20s file that decodes into 62 segments of correct lyrics ungated. The file is now retried once with the gate off when the gated pass leaves near-silence, and the retry is discarded if it is a repetition loop. On a 200-file corpus, files with a transcript went 179 → 194.
  • Search: quoting a line verbatim now uses literal matching. It previously ran only as a fallback when the semantic arm returned *nothing*, so the query shape that needs it most got no help whenever the bi-encoder returned anything.
  • Search: the people channel can be ranked. It returned every appearance unranked and with no score — a name resolving 367 rows — which the score-based cross-channel fuser could not place at all, so it contributed nothing even when the gallery had named the person correctly. Rows now carry a presence score, are ranked and capped, and are anchored to their frame so they corroborate the visual/OCR hit for that instant instead of competing with it.
  • Search: a person named inside a longer query resolves. Matching was one-directional, so "Xi Jinping on the trade war" resolved to nobody while "Xi Jinping" alone resolved to 95 appearances.
  • OOMSimulator.get_peak_rss_mb() could report a peak below the current RSS.
Changed
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current main release head with synced 0.1.43 package, desktop, MCP bundle and storefront metadata.
  • The bundled gallery of public figures gained the names an earlier collection run had quietly dropped — Bill Gates, Marie Curie and Alan Turing among them. Every added portrait's licence was resolved; all free.
Added
  • MEDIAFIND_ASR_VAD, MEDIAFIND_ASR_VAD_FALLBACK, MEDIAFIND_ASR_VAD_MIN_COVERAGE, MEDIAFIND_ASR_VAD_MIN_TTR, MEDIAFIND_TRANSCRIPT_LEXICAL, MEDIAFIND_TRANSCRIPT_LEXICAL_FLOOR, MEDIAFIND_TRANSCRIPT_LEXICAL_K, MEDIAFIND_PEOPLE_K, MEDIAFIND_PEOPLE_PRESENCE_FULL, MEDIAFIND_PEOPLE_PRESENCE_FLOOR (see docs/configuration.md).
0.1.42
Added
  • Suggestions drawn from your own library — the empty search surface now offers the people, organisations, places and actions the index actually found in your media, instead of only saved searches, categories, and three hardcoded example prompts. Every chip is a term the library contains, so it always returns something. Keyless and deterministic — no model runs.
  • Product and technical decks — added a polished intro deck, technical documentation deck, and long-form technical documentation for sharing how MediaFind is built and what it ships.
Changed
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current main release head with synced 0.1.42 package, desktop, MCP bundle and storefront metadata.
Fixed
  • Model loads cannot wedge indexing — every embedding, vision, face, and transcription model load now runs behind a deadline so a stalled fetch cannot leave the indexer stuck forever.
  • Bug-bash reliability fixes — hardened older-library migrations, destructive action exports, object detection floors, brand/logo substring matches, crashed-job recovery, and cleanup deduplication so recovery paths keep a readable copy.
0.1.41
Added
  • Celebrity recognition reaches far more public figures — the bundled reference gallery is much larger than before and keeps everyone it recognized before, widened past politics/film/music/sport into television, journalism, modelling, more sports, chefs, and astronauts. Still keyless and on-device: only derived numeric descriptors ship, never images.
  • One naming rule for your labels — tags, collections, and saved searches now share a single normalization: "Work" and "work" can no longer become two separate tags that lookups silently blend, every category (including multi-word ones like "screens & documents") is reachable from the search box, category names match in search results, and mediafind doctor reports any legacy case-duplicate labels without touching them.
  • All labels in one request — a new /api/media/labels endpoint returns a file's categories, tags, collections, chapters, notes count, and saved- search hits together.
  • Library chapters → Create projects — imported in one click-shaped API: a project can pull its source video's auto-generated chapters into the editor timeline (exact frame math, NTSC-safe).
  • Entity mentions link to people — anchored mentions in the Knowledge entities rail jump straight to that person's identity in Faces & People.
  • Identities in Faces & People — the canonical people roster is now visible and editable: each identity shows its face/voice/mention/gallery links, with rename, merge, and per-link detach right in the panel.
  • Stale-citation warnings — reopening an Ask conversation after files changed now flags affected citations ("source changed", "file moved", "file missing") instead of silently presenting outdated quotes as current.
  • Storage view in Cleanup — the data folder's per-directory usage and a one-click orphaned-file sweep (dry-run first, confirmation required) live in the Cleanup surface.
  • One person, one node — the Knowledge map now joins faces, voices, and transcript mentions on canonical identities: a person you've linked across facets appears as a single node with all their files, and entity mentions anchor to a stable identity that survives re-clustering and renames.
  • Identity-tagged search results — results carrying a recognized speaker, person, or entity mention now include the canonical identity, so grouping "everything with this person" across channels is possible.
  • People identities — a person now has one stable identity across face clusters and voices: confirming a voice↔face link mints a canonical identity that survives reclustering, renames, and merges, exposed via a new /api/identities surface (roster, merge, rename, attach/detach).
  • Ask answers you can trust later — every Ask answer and its citations are now snapshotted (quote text, hash, and the ASR/embedding/LLM models that produced them). Ask history survives an app restart, and a new verify endpoint reports when a cited source has since been re-transcribed, moved, or deleted instead of silently showing changed text as current.
  • Storage report & cleanup sweep — see exactly what MediaFind's data folder holds (thumbnails, clips, waveforms, reels, exports, models, …) and reclaim orphaned derived files, from Settings-level API (/api/cleanup/storage), or the new mediafind storage CLI command.
Changed
  • Knowledge rail loads leaner — the entities rail now carries its person links in the page data it already fetches, so browsing Knowledge no longer makes an extra round-trip on every render.
  • Engine restructure (cont.) — the media-ingest write path moved into its own module; the index engine is now 54% smaller than where this refactor started, split across focused modules instead of one file.
  • Frontend restructure — the page JavaScript moved out of the single large HTML template into ten focused modules under static/js/, and the search monolith split into five more; the served page is byte-identical, but the code is finally navigable.
  • Engine restructure — the index engine's schema/migrations, detection facets, and ANN machinery now live in their own modules, four dead internal facade layers were removed, and a dozen web routes moved out of the app composition file into their topical route modules.
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current main release head with synced 0.1.41 package, desktop, MCP bundle and storefront metadata.
Fixed
  • Older libraries link people everywhere on first open — a library indexed before the identity layer now upgrades its entity mentions to canonical people as soon as it is opened, instead of waiting until the entities view happened to be visited.
  • Re-clustering no longer strands links — canonical identities and link-review decisions now follow people through a full face re-cluster even when the person was named but never voice-linked.
0.1.40
Added
  • Broader release benchmarks — added public-dataset benchmark coverage for the previously unmeasured search channels so release quality signals cover more of the shipped retrieval surface.
Changed
  • Object-search configuration — made config.py the single home for object tunables, reducing drift between preferences, runtime behavior, and benchmark setup.
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current main release head with synced 0.1.40 package, desktop, MCP bundle and storefront metadata.
Fixed
  • CI latency gate under fleet contention — stopped transient runner contention from blocking the evaluation latency gate during release validation.
0.1.39
Fixed
  • MCP server dropped its connection mid-session — under the stdio transport stdout *is* the JSON-RPC channel, but engine diagnostics were written there with bare print(). A buffer flush spliced raw text into a reply, which the client could no longer parse; it surfaced as an unexplained timeout or "server disconnected" long after the write that caused it. Stdout is now reserved for the protocol and stray writes go to stderr.
  • MCP servers failed to start in GUI clients — Claude Desktop, Cursor, Windsurf and Zed launch with a bare login PATH, so a mediafind-mcp command that works in a terminal was never found by the app. Configs now name an absolute interpreter.
  • Claude Desktop bundle — carried the same PATH failure, and its version had drifted 17 releases behind because nothing rebuilds it on a release. Both fixed, and make sync now guards the manifest version.
Added
  • mediafind-mcp --print-config <client> — emits a ready-to-paste MCP config for Claude Code, Claude Desktop, Codex, Cursor, Gemini CLI, VS Code, Windsurf and Zed, in each client's own format with a PATH-independent command.
  • mediafind-mcp --doctor — diagnoses an MCP connection by speaking the protocol to it: real handshake, tool calls, stdout-purity check and timings. It distinguishes a corrupted stream from cold-start latency, which matters because corruption presents as a timeout.
  • MediaFind agent skill (packaging/skill/mediafind-mcp/) — installs and verifies the MCP server across clients.
Changed
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current main release head with synced 0.1.39 package, desktop, MCP bundle and storefront metadata.
0.1.38
Added
  • Developer Pro unlock — hidden development builds can now unlock Pro locally with the guarded dev-mode environment flags, while production builds continue to ignore local-only unlocks.
Changed
  • Evaluation campaign durability — hardened exact public-benchmark and supervised model campaign tooling with disk/headroom guards, source-group validation, query grouping, partial materialization support, and clearer readiness provenance.
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current main release head with synced 0.1.38 package, desktop, and storefront metadata.
0.1.37
Changed
  • Shared license activation — cut over the desktop app to the shared licensing service so current builds can activate successfully against the production license backend.
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current release head with synced 0.1.37 package, desktop, and storefront metadata.
0.1.36
Added
  • Public-video training pipeline — added receipted source-audit, materialization, feature-building, checkpoint, and leaderboard-readiness tooling for the public-video-benchmark top-contender campaign.
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current release head with synced 0.1.36 package, desktop, and storefront metadata.
Changed
  • Evaluation readiness — added disk/headroom preflights, bounded visual feature caching, cache-local media grouping, and clearer readiness ledgers for partial exact campaign runs.
Fixed
  • Search channel restoration — restored the temporal action search channel and corrected OCR sidecar card assignment in result rendering.
0.1.35
Added
  • ShadowIndex — infer and search for missing media that is not in the current library using corroborated evidence from existing files, with an opt-in Pro-gated search channel, scan job, review API, and evidence-rich UI cards.
  • Exact public video-eval runners — added public-benchmark video-evaluation tooling, submission prep, aggregation, and model-backend coverage for reproducible retrieval benchmarking.
  • LAN access toggle — let users explicitly expose the local MediaFind server on the LAN so they can search from a phone or another device on the same network.
  • Launch materials — added the Product Hunt launch kit and pinned the Android companion APK link for the storefront.
Changed
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current release head with synced 0.1.35 package, desktop, and storefront metadata.
Fixed
  • Bug-bash fixes — shipped the July 17 licensing, search, Create-room, downloader, frontend, and server-guard fixes verified across the release suite.
  • Release-gate hardening — bounded long-running benchmark campaign subprocesses and made the batch download/index job test tolerate slower loaded release machines.
0.1.34
Changed
  • Cross-platform release refresh — rebuilt the signed, notarized macOS DMG and native Windows/Linux installers from the unchanged v0.1.33 source head, with synchronized 0.1.34 package and desktop metadata.
  • Storefront release alignment — advanced the public download URLs and static asset cache keys together so the website, GitHub Latest release, and in-app stable update channel all resolve to v0.1.34.
0.1.33
Added
  • Confidence-aware multimodal fusion — combine transcript, visual, OCR, and metadata evidence with calibrated channel confidence and learned local weights so mixed-signal searches rank the strongest moments more reliably.
  • Temporal action recognition — recognize actions across frame sequences with a bundled action model over cached visual embeddings, without uploading media or re-encoding the source library.
  • Auditable Perception Test benchmark — reproduce temporal-grounding results with pinned attribution, metric protocols, evaluator provenance, and official-union IoU parity.
Changed
  • Release quality gates — expanded regression coverage for fusion and Perception Test evaluation, including refreshed baselines and model-portfolio documentation.
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current release head with synced 0.1.33 package, desktop, and storefront metadata.
Fixed
  • Create-room export destinations — use the saved-path fallback consistently for FCPXML, EDL, and caption exports when a native save dialog is unavailable.
  • Premiere timeline timecode — carry fractional seconds into the next minute correctly instead of displaying values such as 59.96 seconds.
0.1.32
Added
  • Private Media Intelligence workflows — turn local transcript and visual signals into cited Interview Briefs, Research Quote Packs, and Rough Cuts, with natural-language event queries, reviewable speaker-to-face suggestions, and an explicitly opt-in local visual verifier.
Changed
  • Subscription licensing — moved MediaFind Pro to monthly and annual Stripe subscriptions with signed device-bound leases, self-service billing, and controlled device transfers.
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current release head with synced 0.1.32 package, desktop, and storefront metadata.
Fixed
  • Cross-platform library paths — display folder and file names correctly when a saved library contains Windows, Linux, or macOS path separators.
  • Notes and export durability — preserve note ranges and colors during partial edits and imports, avoid concurrent export clobbering, and harden cleanup of malformed note data.
0.1.31
Added
  • Moment DNA reattachment — recover annotations after media is transcoded, reframed, or otherwise transformed by aligning stored visual moments, with a review panel on the file page for applying recovered matches.
  • Correction Memory — remember approved search corrections locally and surface recovered results and counts in the interface so a correction made once can improve future retrieval without sending library data off-device.
Changed
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current release head with synced 0.1.31 package, desktop, and storefront metadata.
Fixed
  • Text export destinations — added a save-to-folder fallback for coded notes, calendar, and timeline text exports when the native save dialog is not available.
0.1.30
Added
  • More accurate scene recognition — bundled calibrated on-device classifiers for eight broad scene types and 35 specific places, substantially improving coverage and precision without re-encoding media or sending it off-device.
Changed
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current release head with synced 0.1.30 package, desktop, and storefront metadata.
Fixed
  • Scene search precision — required whole-word matches for curated scene labels so unrelated fragments such as "ice" or "officer" no longer activate the Office facet.
0.1.29
Added
  • Local-first library sync — synchronize MediaFind libraries between devices through a user-selected shared folder while keeping the source media and index under the user's control.
  • Production search vocabulary — search and filter by shot type, camera angle, camera motion, HDR, bit depth, color space, B-roll, cuts, montage structure, alternate takes, and resurfaced people-and-date memories.
  • Premiere Pro panel — search the MediaFind library and add results to the active timeline from a live in-editor CEP panel.
Changed
  • Faster, more resilient indexing — parallelized folder and pipeline stages, auto-tuned concurrency, bounded nested frame workers, and healed renamed or moved files without unnecessarily rebuilding their index records.
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current release head with synced 0.1.29 package, desktop, and storefront metadata.
Fixed
  • Search and metadata correctness — refreshed knowledge caches after speaker or entity changes, folded diacritics consistently, preserved same-name local collections, and bounded malformed embedding slices and clip values.
  • Media workflow reliability — corrected diarization scoring, atomic share bundles, OTIO file URLs, C2PA manifests, music-gap detection, and settings responsiveness during indexing.
0.1.28
Added
  • Cross-file Ask — added a conversational across-files mode that synthesizes answers by source while preserving follow-ups, suggestions, and source links.
  • Report exports — surfaced clean, timestamped transcript downloads and CSV quote-table export for citations gathered across multiple files.
Changed
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current release head with synced 0.1.28 package, desktop, and storefront metadata.
  • Marketing-site polish — repaired navigation toggles, product-page check marks, long-title hero layouts, tutorial overlap, and narrow reading columns.
Fixed
  • Request hardening — bounded cleanup, restore, purge, forget, and protected directory lists and their path lengths without constraining folder removal.
  • Ask source navigation — kept repeated quotes from different files linked to their own source cards in cross-file answers.
0.1.27
Added
  • Obsidian export — exposed an Obsidian-friendly export path alongside the hardened export, reel, and clip-pack surfaces.
  • Storefront download metrics — activated the get.mediafind.io download counter across MediaFind and sibling apps, with D1-backed asset resolution and a cache fix so redirects stay current across Cloudflare colos.
  • Video benchmark evals — added new benchmark coverage for video-oriented quality and performance checks.
Changed
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current release head with synced 0.1.27 package, desktop, and storefront metadata.
  • Export UX — surfaced redaction status and per-speaker transcripts in the web UI so exported material is easier to review before handoff.
Fixed
  • Security hardening — closed a second symlink-write hole, bounded waits and request fields that gate real work, and tightened export, reel, and clip-pack surfaces.
  • Ask grounding — stopped citing sources for "I don't know" answers and fenced untrusted transcripts in LLM prompts.
0.1.26
Added
  • Campaign attribution — added site-wide Plausible support and UTM passthrough so campaign tags persist to Polar checkout links without adding a backend.
Changed
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current release head with synced 0.1.26 package, desktop, and storefront metadata.
Fixed
  • Bug-bash hardening — fixed correctness and durability issues across search, license handling, cleanup, create/export rendering, provenance, indexing, and the home/onboarding UI.
0.1.25
Added
  • Insights discovery — added Trending Topics and Sound & Music rails, made user notes searchable by default, and let users add more folders/files while indexing is still running.
Changed
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current release head with synced 0.1.25 package, desktop, and storefront metadata.
  • Knowledge graph clarity — explained what graph circles mean and tuned people/voice suggestions so "View clips" routes to the right surfaces.
  • Storefront refresh — bumped the PhotoFind download page to its 0.1.1 release.
Fixed
  • Insights browsing — made "Most of your media is from <month>" Browse open the matching media instead of a dead-end filter.
  • Search and indexing hardening — fixed bare codec-token filtering, inherited interactive busy timeouts across channel fan-out, and serialized refresh-faces detection against the native index lock; speaker-filtered searches no longer leak file-level Details hits.
  • UI polish — fixed left-aligned chapter rows, one-letter follow-up columns, the "Your tags" page chip markup, graph drag/pan release navigation, everyday gallery mononym extraction, download job terminal copy, and the meeting-processing button label.
0.1.24
Changed
  • Cross-platform release — refreshed the signed, notarized macOS DMG and native Windows/Linux installers from the current release head with synced 0.1.24 package, desktop, and storefront metadata.
  • Site expansion — added a Descript-alternative comparison page for people evaluating private, local media search alongside cloud editor workflows.
Fixed
  • Bug-bash hardening — fixed correctness, privacy, and durability issues across imported-media trust state, sequential-audio filtergraphs, scoped search recall, stale frame ANN signatures, malformed captions, GIF time caps, audiogram and waveform edge cases, CLI safety checks, object detection, phonetic matching, provenance revocation, phone redaction, and atomic saved-search refresh.
0.1.23
Added
  • Task Assistant — added the draft-only goal-to-recipe router, approval-gated render-reel action, gated batch tagging, and duplicate-review workflow so assistant-suggested library actions stay explicit and reviewable.
Changed
  • Site navigation — completed the Apps and Developers navigation rollout across the marketing pages and added SDK/MCP education for people embedding MediaFind locally.
Fixed
  • Bug-bash hardening — fixed correctness, privacy, and security issues across share redaction/source handling, unsupported password-protection paths, missing-transcript state, color facets, cancellation races, malformed face blobs, bare-word source filters, symlink-safe import/export/download paths, mixed-fps clip-pack trims, Ask anti-repeat grounding, ANN atomic swaps, and search-pool load shedding.
  • Release verification — fixed release-status false negatives so the post-publish verifier reports half-published releases more accurately.
0.1.22
Added
  • AI-agent integration — added a local MCP server so Claude Desktop, Claude Code, Cursor, and other MCP clients can search, ask, and read transcripts from the same on-device MediaFind library.
  • In-app agent setup — added a Settings flow that shows copyable MCP commands/JSON and can install the Claude Desktop config when the local desktop environment allows it.
  • Python SDK — introduced the typed mediafind.core.Engine facade, exported it from mediafind, and shipped a py.typed marker for downstream type-checkers.
  • Public HTTP API — added the frozen-contract /api/v1 routes for search, Ask, transcript reads, and status so local and on-prem integrations do not depend on internal UI endpoints.
  • Developer and agent docs — added the public Developers and AI agents pages, plus MCP packaging docs and the MediaFind MCP product icon.
Changed
  • Cross-platform release — refreshed the signed, notarized macOS DMG and published native Windows and Linux installers from the current release head with synced 0.1.22 version metadata.
  • MCP privacy controls — added operator control for whether MCP tools may return answer and transcript content to the calling agent.
  • Website updates — surfaced the AI-agent and developer surfaces from the home page and Tools menu, added the PhotoFind product page to the site family, and added new comparison/intent guide pages.
Fixed
  • Packaged startup checks — avoided unnecessary model probes for empty-library status, health, and stale-index checks so a fresh packaged app starts without loading local AI models just to report idle state.
  • Release QA stability — added a server-only packaged-app QA mode and longer live CUJ timeout so signed release verification exercises the app without window-management flakiness.
0.1.21
Fixed
  • First-run transcription stays offline — packaged builds now use the bundled transcription model on first use instead of falling back into a runtime model download.
  • Long-file diarization stability — bounded diarization work on long files and made speaker-stage progress visible during indexing.
  • Indexing throughput — background jobs now run at the expected system QoS so indexing work uses performance cores again.
  • Packaged Ask responsiveness — bounded bundled Mini answer generation so a cold packaged-app Ask request cannot monopolize the release CUJ run.
  • Ask cold-start fallback — default Ask now returns a cited extractive answer until the local model is already warm, avoiding cold-start stalls.
  • Summary cold-start fallback — default indexing no longer starts the bundled local model just to generate summaries; explicit LLM summaries still opt in.
0.1.20
Fixed
  • Indexing freeze hardening — bounded the remaining client-reachable ffmpeg/ffprobe call sites, serialized native media work more defensively, and kept foreground UI reads responsive while a library is indexing.
  • Crash-loop resilience — set the native OpenMP duplicate-runtime guard at package import time and excluded the app data directory from media walks to avoid flash-quits and self-indexing loops.
0.1.18
Fixed
  • External checkout links — the Tauri desktop shell now opens non-local links, including Pro checkout links, in the OS browser instead of trapping them inside the app window.
  • Pro pricing consistency — synced the in-app Pro price constant to the current $29 storefront price.
0.1.17
Changed
  • Storefront cleanup — removed the retired MediaCreate, MediaGen, and MediaTrends marketing surfaces so the site focuses on MediaFind.
  • Pro pricing refresh — updated the site and related docs to the current $29 Pro price and kept purchase links visible during the free trial flow.
0.1.16
Fixed
  • Freeze-prone workflows — hardened long-running create/render work, desktop shell health checks, stale database recovery, and background routing so the app is less likely to stall under heavy local media work.
Changed
  • What's New navigation — moved the What's New callout into the main navigation CTA area so release updates stay visible without crowding the nav.
0.1.15
Changed
  • Offline first-run transcription — packaged builds now bundle the lightweight Base transcription model, so new users can transcribe local media without a first-run ASR download; Small and larger models remain quality upgrades.
0.1.14
Fixed
  • Search no longer freezes during indexing — heavy media work (transcription, embeddings, frames, faces) now runs through the background job queue instead of blocking the search path, so the app stays responsive while a library indexes.
  • On-screen-text (OCR) matches — an exact on-screen-text hit is now treated as a confident match in search results rather than being scored down.
  • First-run model picker — improved the default model selections offered on first run.
  • Stabilization guardrails — additional guardrails across background work and UI state to prevent stuck/stalled states, plus a round of release bug-bash hardening and clearing a stale ASR download job id.
0.1.13
Added
  • Per-file Re-transcribe — the play page now has a per-file Re-transcribe button to re-run ASR on a single recording on demand.
  • ASR download status in Settings — Settings now surfaces the current ASR model download status.
Fixed
  • Apple GPU engine empty transcripts (frozen app) — the Apple GPU engine now decodes audio in-process instead of shelling out to a PATH ffmpeg that the packaged app lacks, so transcripts are no longer silently blank; also fixed the Apple GPU engine's model directory.
  • Stranded empty transcripts — when the ASR model isn't ready at first index the file is no longer left permanently blank; indexing logs a warning and self-heals on a later pass.
  • Search freezes during indexing — hardened the search, playback, and live cold-load paths so the UI no longer freezes or times out while indexing is running.
  • Frozen / stuck UI states — audited and hardened remaining stuck job and UI states, the activity-dock close button, model-change settings UX, and queued exports / architecture lifecycle.
0.1.12
Changed
  • Keyless AI backend — removed the old keyless answer/summary backend and made the bundled Mini on-device model the default. Ask and summaries now work keyless, on-device, with no cloud, account, or download out of the box; Small is the one-tap quality upgrade.
  • Unified Activity dock — all progress and status (indexing, downloads, model fetches, background jobs) now live in one always-present Activity dock, replacing the per-feature progress bars and the old global job tray.
Added
  • Analysis sidecar import — indexing now reuses matching, non-redacted <media>.mediafind.json sidecars to restore transcript segments, summaries, and chapters before recomputing expensive analysis.
  • MediaCreate — an on-device video studio (timeline editing, transitions, color, reframe, titles, audio, render presets) fed straight from search results, with FCPXML/EDL/OTIO/AAF interchange export and OTIO re-import.
  • Celebrity recognition — expanded the bundled gallery of public figures and raised the match threshold for more precise name matches.
Fixed
  • Indexing correctness — hardened the local and internet indexing paths.
  • Visual frame alignment — a frame whose processing failed no longer shifts every later frame onto the wrong visual embedding, timestamp, and dominant color.
  • LLM chapter titles — chapter titles now use the bundled on-device model before falling back to a local LLM server, matching summaries.
  • Cleaner AI summaries — non-speech ASR markers (e.g. blank-audio / silence / inaudible) are stripped before summarization, so a near-silent clip yields an empty summary instead of a summary of those markers.
  • Job metrics — failed and cancelled jobs are no longer counted as completed in the Activity dock totals.
  • Search paywall — a People & Faces search no longer shows a "Brand-logo" upgrade card; the paywall is titled from the actual gated feature.
  • Face-indexing defaults — the Add Media face checkbox is now selected by default only for Pro/trial installs, and left unchecked for free installs, matching the privacy copy and product gating.
  • Faces never stranded — a file is no longer marked faces-done when the face backend was unavailable; indexing retries instead of silently leaving zero faces.
  • Download start — fixed a freeze when starting a download, and the activity-dock Pro badge no longer overlaps adjacent UI.
0.1.11
Added
  • Meetings mode polish — a batch-first Meetings redesign with cross-meeting follow-up rollups, a Source/transcript toggle, Copy-as-Markdown, a bundled policy-meeting demo clip, and a Meetings walkthrough in the first-run tutorial. Curated action items and decisions load straight from a meeting's summary.json sidecar — now honored even for already-indexed recordings.
  • Live indexing progress — adding videos now shows a single, expandable job card with named progress stages instead of a flashing, multi-card status.
  • Faces selected by default when adding media — face indexing is pre-checked on the Add Media page (still Pro-gated; it skips silently for free users).
Changed
  • Category chips open in Home — clicking a category now routes to a precise in-Home browse view rather than jumping to Find.
  • Faster app launch — the desktop window now opens immediately and navigates to the UI as soon as the local server is ready.
  • Sweeping accessibility pass — hundreds of buttons, chips, toggles, tabs, and dialogs across Find, Ask, Knowledge & Insights, People & Speakers, Collections, Cleanup, Operations, Add Media, and the Video Editor now carry proper labels, hover titles, ARIA roles/state, keyboard focus handling, and screen-reader announcements for empty-state and result feedback.
  • Clearer empty-input feedback — submitting an empty search, Ask prompt, tag, collection, or saved search now announces what's missing instead of failing silently.
  • Reliability and performance hardening across the API — async request handling for heavy routes, a thread-safe search-index connection pool, opt-in structured request logging, a feature-flag system, and a unified error-response format — with no change to on-device, keyless behavior.
Fixed
  • Search reliability under load — the shared search-index connection pool now opens its SQLite connections for safe cross-thread reuse, fixing intermittent 500s (e.g. on the suggestions rail) when the server handled concurrent requests.
  • Offline faces — face-recognition weights are now bundled, so face detection works in the packaged app without a network round-trip.
  • Onboarding & status routes — fixed an admin-router prefix bug that broke the onboarding and status endpoints.
  • Model picker — the local-LLM picker now unlocks when a local LLM server is running, keeps the local model choice available, and focuses the currently selected model.
  • Insights — findings drill-through now handles the "open" action type, and the Embeddings coverage meter uses the transcribed-file count as its denominator.
  • Home — face crops now appear in the asset quick-look drawer's People section.
  • Sample clips — corrected the playback allowlist and face-thumbnail fallbacks so bundled demo clips and player faces render reliably.
  • Library imports — imports are now cancellable, abort cleanly if a safety backup fails, preview their impact before overwriting (dry run), and clean up pending uploads on cancel.
  • Downloads — batch download jobs now fail correctly when every item fails, surface backend hints in discovery, and re-enable cancel for batches.
  • Captions — generated caption tracks are replaced on regeneration, remapped to sequence time (including translated and speed-adjusted clips), and sorted by sequence time.
  • File resolution — segments, notes, tags, transcript edits, frames, and exports now resolve media by unique basename, with ambiguous CLI basenames rejected instead of guessed.
  • Privacy — share output paths stay hidden by default and the privacy audit keeps running on keyless backends.
  • Resolved a large batch of additional code-audit findings across search, migrations, exports, provenance, and meetings endpoints.
0.1.10
Added
  • Create — an on-device video studio. Turn search results into a finished video without leaving MediaFind: send clips straight from Find to a new project (Find → Create handoff), then edit on a real timeline with split, drag-trim, drag-move, duplicate, reverse, speed changes, rotate/flip, and undo.
  • Transitions & motion — cross-dissolve and fade-to/from-black between clips, keyframed overlay animation, and freeze-frame holds.
  • Color room — named look presets and primary color adjustments.
  • Reframe room — re-aspect to any ratio with subject-tracked cropping, previewed live (WYSIWYG).
  • Text & Graphics room — titles and lower-thirds with fade/slide-in animation and brand-kit styling.
  • Audio room — per-track gain, loudness normalization, denoise + high/low-pass EQ, dissolve crossfades, and auto-ducking (music under speech).
  • Brand Kits & Templates — save reusable brand styling and whole-project templates, and start new projects from a template.
  • WYSIWYG preview — real frame-at-playhead video preview with live caption/graphic overlays, reframe crop, and color grade applied.
  • Deliver — render presets (resolution / codec / container) with proxy-generation control.
  • More export formats — animated GIF, contact-sheet thumbnail grids, audio-only export, poster/freeze frames, project chapters (YouTube / ffmetadata), captions (SRT/VTT), and full project JSON export/import portability.
  • Multi-language caption translation.
  • Content provenance — C2PA Content Credential badges on Find results.
  • Apple GPU / Neural Engine transcription — a GPU/Neural-Engine transcription engine now ships in builds for faster on-device transcription.
  • Durable crash logging — on-disk logs plus crash handlers for better support diagnostics.
Changed
  • Hardened security and reliability across the API — path-traversal, SSRF, SQL-injection, and import-size protections — plus new per-channel search-latency and queue-depth metrics.
Fixed
  • Share bundles — output paths are now redacted by default; share IDs no longer accept fake/spoofed values.
  • Face detection — reverted to opt-in default (was accidentally flipped on in 0.1.9).
  • Search latency metrics — Prometheus channel-latency histogram now correctly selects the child label before observing.
  • Job queue — hardened handling of closed and legacy job states to prevent spurious errors.
  • Collections — rename errors now surface context instead of a blank message.
  • MediaCreate — undo and transitions repaired; add-clip is now undoable.
  • Exports — library exports written atomically; prior render exports preserved on failure; report exports no longer clobbered.
  • Interchange — default FPS validation added; render container validation tightened.
  • Reindex — stale-embedding detection now triggers when sidecar inputs change.
  • Embedding model — model ID stays tied to the loaded backend across restarts.
  • Resolved additional code-audit findings across packaging, API validation, and meetings endpoints.
0.1.9
Added
  • Unified Find + Ask surface — search and the conversational agent now live on one Google-style surface instead of separate destinations.
  • On-device song recognition — keyless music detection plus same-track grouping by chroma fingerprint, with a player "Same music track" panel that lists every clip sharing the recognized track. Fully on-device; no AcoustID.
  • Faces on by default — face detection is now enabled by default everywhere (still Pro-gated: it skips silently for free users and 402s on explicit opt-in).
  • Multilingual search — an opt-in i18n embedding profile for non-English media, selectable from a new Settings → Search language / Embedding model toggle.
  • Knowledge & Insights dashboard — a chart dashboard and a capture-date timeline over your library's facets.
Fixed
  • Packaging — ship the runtime resource directories inside the wheel so a pip install has everything it needs at runtime.
  • Dependenciesrequirements.txt is back in sync with the pyproject core deps.
  • Auto-updater — replace the deprecated datetime.utcnow() with a naive-UTC helper.
  • Play page — focused person page and scroll-position restore when going Back from a facet.
  • Resolved 23 confirmed code-audit findings.
Changed
  • Marketing site — light/dark theme switch (light by default) with theme-aware blog illustrations.
  • Internal — typed domain errors with a structured handler across the API, a continued app.pyapi/routes/ decomposition, index.html inline JS split into static modules, and a per-commit retrieval quality + latency eval harness with a regression gate.
0.1.8
Added
  • NLE interchange export — send search moments to a video editor as an FCPXML/EDL timeline for Premiere, Resolve, or FCP (POST /api/export/timeline, Pro). Each clip is marked with its matched text, exports use drop-frame timecode for 29.97/59.94 footage, and a "bundle trimmed clips" mode produces a self-contained proxy bundle. A one-click handoff sends found moments straight to the open DaVinci Resolve timeline (POST /api/export/resolve) via a bundled, installable Resolve script.
  • Nine new keyless search channels — scene, audio, object, color, emotion, phonetic, temporal-action, related, and by-image, alongside a channel/scope picker in a unified search bar.
  • Named-entity facet — exact-name search over transcripts and on-screen text via a keyless gazetteer, with an optional open-vocabulary NER backend and Wikidata linking.
  • On-screen text (OCR) is now its own selectable search modality.
  • Brand-logo and action detection facets — keyless zero-shot visual search for brand logos (by name or sample image) and actions (e.g. "dancing", "cooking").
  • Ask, redesigned as a multi-turn conversational agent and promoted to its own top-level destination, with multi-select scope (limit answers to chosen folders/files) and per-bubble citations.
  • Knowledge & Insights — the Knowledge Agent reworked into a facet-aware dashboard with coverage meters, on-device AI summaries, and a findings inbox.
  • Transcription model picker — choose a transcription model on first use, with an optional Metal engine and persisted word timestamps + confidence.
  • On-device LLM tier — download a local model from Settings to power Ask and summaries fully offline.
  • Video Editor — a keyless ffmpeg Tools workspace (cut, convert, resize, compress, extract audio, grab frame, GIF) with a live preview and trim timeline.
  • First-launch guided tour — a spotlight walkthrough of search, Ask, People, and brands over the demo clips, re-runnable from Settings.
  • Demo mode and a refreshed onboarding sample set whose bundled clips exercise every demonstrable search channel.
  • Player — pop video out to a floating/Picture-in-Picture mini player, draggable anywhere, with finer waveform scrubbing.
  • People & Voices — merge/link voices, per-person appearance segments + clip export, click-through face/voice/brand/action pages, and celebrity detection.
  • Search — folder and date-range filters, recency sort, and bulk-select on results and the library grid.
  • Mobile — the companion app codebase is now cross-platform (iOS + Android) with OTA updates and a no-server demo mode; Android ships as a sideloadable APK, with iOS via TestFlight planned.
  • Reverse-image search is now a first-class channel (no longer logo-gated): find visually similar frames from any sample image.
  • Comprehensive file detail on the play page — full metadata, categories, brands, actions, and an on-screen-text (OCR) panel that now also covers images.
  • Richer onboarding demos — an OCR-showcase clip, plus curated summaries and chapters for the bundled demo clips.
  • Video Editor — AV1 output in the convert tool.
  • First-run model pickers block until the download finishes, showing progress, so Ask/summaries are never silently downgraded; the Pro trial counts down in hours on its final day.
Changed
  • Large-library scaling — heavy operations moved off the request thread, hot aggregations cached, and writes batched; load-shedding sheds search overload (503 + Retry-After) while keeping /health responsive.
  • Persistent writable preferences with open-data-folder and clear-cache actions in Settings.
Fixed
  • Resolved findings from multiple full-codebase, UX, and critical-user-journey audits — relevance floors so off-topic queries don't surface junk, honest done/empty states, accessibility and contrast fixes, and several latent crash paths in faces, jobs, and people.
  • Cross-platform: persistent data directory and UTF-8 file IO on Windows/Linux.
0.1.7
Added
  • Click-a-voice to see a speaker's clips, mirroring click-a-face; player panel scoped to the video with voice management moved into Faces & People.
  • Right-click a library file to reveal it in Finder, play it, or copy its path.
Fixed
  • Scroll-away player keeps result-navigation context, and face-click playback jumps to the right moment.
0.1.6
Added
  • Cross-platform desktop installers: native Ubuntu + Windows builds via Tauri, alongside the macOS app, plus matching Windows/Linux (beta) downloads on the site.
  • "Try MediaFind on real videos" onboarding: three bundled, redistributable demo clips that install + index offline (GET /api/samples, POST /api/samples/install).
  • Dedicated player page reused across search, play, and full-view open.
  • In-app blog with three technical deep-dives.
Changed
  • Cosmetic overhaul (continued): a single SVG icon system across app chrome, sidebar nav, and result rows, plus an inline-style teardown into named components (Dialog) and utilities. Unified Pro-upsell CTA wording.
Fixed
  • Accessibility: stop exposing a phantom second search field to screen readers; restore focus to the trigger when the image lightbox closes.
  • Show an inline error + Retry instead of silently blanking panels on failure.
0.1.5
Added
  • Duplicate & cleanup workspace: near-duplicate / similar-video detection with a safe, reviewable quarantine workflow, batch keep-best policies, per-frame fingerprinting, and quality culling of blurry / low-res images (/api/cleanup/* — exact-dup free, near-dup / quality tiers Pro).
  • Per-person video segments + clip export: click a person to see contiguous appearance spans and export them as clips or a stitched reel (GET /api/people/segments, POST /api/people/clips).
  • In-app feedback: report bugs, send thanks, or suggest ideas, logged locally and sent via the OS mail client (/api/feedback).
  • Finder-style folder tree + resizable rows in the library list view.
  • Unsigned Ubuntu + Windows server binaries, plus an enriched commercial landing site.
Changed
  • Cosmetic overhaul (phase 1–2): a design-token + utility foundation and an explicit button-tier system.
  • Honest "Refresh embeddings" nudge surfaced inline when search is degraded.
  • Mark Pro-gated sidebar items with a PRO badge for free users.
Fixed
  • Spoken-word search and Ask now keyword-fall-back so they work when embeddings are stale instead of returning nothing.
  • Mouse / browser Back button navigates between workspaces.
  • Knowledge graph canvas drag / zoom / click restored.
0.1.4
Added
  • Redesigned Home / library browser: a dedicated Home page (split from Find) that lays files out as icons or a list, with sorting, folder drill-in, and an adjustable icon size; a new sidebar of Recent files, Recent searches, and Saved searches.
  • Knowledge Agent: Categories and the Knowledge Map became an evidence-first agent — a proactive review/findings queue, a durable inbox, ad-hoc Q&A, and a live force-directed graph of how the library connects.
  • Celebrity / public-figure face recognition: notable faces are auto-named ⭐ with no setup; click a face to jump to every moment they appear, surfaced in People & Faces.
  • People & Faces rebuild: merge/split people, rich per-person appearances, reindex faces from the page, and automatic clustering after a download.
  • Media downloader: discover videos on a webpage and batch-download a whole channel or playlist, then transcribe and index them like local media.
  • Editable speaker names from the transcript and a speaker library, with the roster moved into the file panel.
  • Auto-update check: background poll of the GitHub releases API with a 24h cached result (GET /api/updates), an in-app toast when a newer version is available, and stable/beta release channels (POST /api/updates/channel, env MEDIAFIND_UPDATE_CHANNEL).
  • Library management: relink moved/missing media, review failed files and retry just those, remove/forget a file (incl. on-disk derivatives), and preview an export before applying it.
Changed
  • Add Media panel regrouped into "from local / from internet" plus a Remove Media section; indexing options are selected by default.
  • Pro UX: a Pro badge replaces the buy-Pro CTA, upgrade CTAs are hidden for users who already own Pro, and is-pro is stamped server-side to avoid a CTA flash.
  • Honest, count-accurate category browse with a stronger relevance floor and a scene-tag confidence gate.
Fixed
  • Packaged macOS app: bundle ffmpeg/ffprobe, the face-recognition model, yt-dlp, and the speaker-diarization model so search, faces, the downloader, and diarization work in the frozen app.
  • Knowledge graph repaints on interaction so pan/zoom/drag work.
0.1.0
Added
  • Durable background job queue with SQLite backend (crash-recoverable)
  • SQLite WAL mode, backup/restore API, DB integrity checks
  • Structured JSON logging and support bundle download
  • Security: SSRF redirect validation, CSRF protection, security headers
  • Real-media fixtures and Playwright E2E tests
  • Typed response models and Pydantic v2 schemas
  • ANN vector index (hnswlib) for >10k clip libraries
  • WER accuracy benchmark suite
  • Library export/import (portable ZIP format)
  • First-run onboarding wizard and ffmpeg detection
Fixed
  • Auth token logged in plaintext (now via logger.warning)
  • Inline JS event handlers migrated to addEventListener (CSP prerequisite)