Meetings

Meeting notes with no bot in the call

Most AI note takers work by sending a bot into your call and your audio to their servers. MediaFind does the job on your computer instead. It transcribes as people talk, labels who's speaking, and pulls out action items and decisions in the words people actually used, each linked to the second it was said. Nothing joins the call, and nothing is uploaded.

A cloud note taker is easy to set up and awkward to explain to the room. A participant called “Notetaker” joins, the recording goes to a vendor's servers, and the transcript of your client call, hiring panel or board meeting now lives somewhere you don't control. For plenty of meetings that's fine. For the ones that matter most, it often isn't.

Meetings in MediaFind work the other way around. The audio is captured on your computer and transcribed there by the same on-device Whisper models that index the rest of your library. Rules on your computer write up the action items and decisions, and an optional local model writes a summary. There's no account, no bot and no upload. The meeting then sits in the same private library as your recordings, chats and notes, so you can search and ask across all of it.

Three ways a meeting gets in

What happens while people talk

Live transcription is a chain of small steps, and every one of them runs on your computer:

Audio mic or the call 1 · Find speech 30 ms frames 2 · Transcribe Whisper, on-device 3 · Who spoke named voices Transcript line who · when · words 4 · Rules action items · decisions quoted, with the second Local model AI summary, ~every 45 s labeled model-written Each new line updates the notes while people are still talking. your computer: no bot in the call, no upload
The rules write the action items and decisions, quoting the transcript. The optional local model writes a summary and says so.
  1. Find the speech. Audio arrives at 16 kHz and a voice-activity detector reads it in 30 ms frames. A line ends after 0.7 seconds of silence, or at 20 seconds, so a monologue still produces captions.
  2. Transcribe it. Each line goes to the same Whisper engine that transcribes your files. While someone is mid-sentence, a rougher caption updates every second and a half.
  3. Work out who spoke. Each voice is compared with the voices heard so far and labeled Speaker 1, Speaker 2 and so on. Rename one and the name applies to every line, past and future. If MediaFind holds a voiceprint for that voice, it remembers the name, and in your next meeting that person is recognized.
  4. Update the notes. Every new line re-runs the extraction described below, so action items and decisions appear as they're said.

Two settings improve the transcript before the meeting starts. Settings → Names & terms takes the people, products and jargon your meetings use, one per line. Every file and live meeting is transcribed with them in mind, so a colleague's name comes out as written rather than as a sound-alike. About fifty short terms fit, the first ones win, and a change reaches the next line of a meeting that's already running. The live panel also lets you pin the spoken language. Whisper otherwise re-detects it on every short line, and a two-second line can come back in the wrong language.

If a local language model is available (MediaFind comes with a small one), an AI summary streams in as well, refreshed about every 45 seconds. Beside it, a visual summary shows the topics, each speaker's talk time and the key terms. The summary is labeled for what it is: “Model-written from the transcript — it can get details such as dates or owners wrong.” That's why the notes underneath it aren't written by a model at all.

Action items are quoted, not written

MediaFind finds action items and decisions with rules, not a language model. It reads the transcript sentence by sentence for the phrases people use when they take something on, such as “I'll…”, “we need to…”, “can you…” and “let's…”, and for the ones they use when something is settled: “we decided”, “we agreed”, “let's go with”. What comes out is the sentence itself, word for word. The owner is the person who said it, or the person it was addressed to by name. The due date is the phrase that was spoken, and every item links to the second it was said.

Here's what the current release makes of six lines from a meeting recorded on Thursday, 13 November 2025:

Said in the meetingWhat MediaFind makes of it
Priya: “Sam, can you send the revised budget by Friday?”Action item · owner Sam · due Fri 14 Nov
Sam: “I'll book the venue next week.”Action item · owner Sam · due Thu 20 Nov
Priya: “We need to close the hiring plan by the end of the month.”Action item · owner Priya · due Sun 30 Nov
Dana: “Let's ship the mobile app by Q3.”Action item · owner Dana · no date, with “by Q3” kept
Priya: “Bob will handle the press release.”Nothing: a promise made in the third person slips past the rules
Sam: “We agreed to go with the river venue.”Decision · Sam

That design gives up some cleverness on purpose. A rule can't invent a task nobody said. It gives the same answer every time, it needs no model and no key, and it's fast enough to run on every line of a live meeting. Because every item carries its timestamp, the words are one click from the moment they were said.

The trade-offs are real, and the table shows one of them: third-person promises get missed. The cue phrases are English, though transcription covers many languages. The rules also assume they're reading a meeting. Run them on a keynote or a film and you'll get a few “action items” nobody meant. The one part of a meeting's brief that a model writes is the Overview at the top.

“By Friday” means the Friday after the meeting

A due date spoken in a meeting is relative to that meeting. So MediaFind resolves it against the date the recording was made. It doesn't use today's date, or the day you got around to processing the file. In the table above, “by Friday” lands on 14 November 2025. Process that recording today and the item shows up as overdue, which is the truth, rather than as due this coming Friday.

That's what makes Follow-ups trustworthy. It collects the open action items from every meeting, grouped by owner, by meeting or by due date, and a second tab lists the decisions. The due-date view sorts items into Overdue, Due this week, Later and No date. A phrase that doesn't pin a day, like “by Q3”, lands in No date with its words kept, not in a bucket it doesn't belong to. Tick an item done (there's an Undo), and export the open ones as an .ics file of tasks.

Across meetings

A meeting isn't a silo; it's a recording in your library. Find searches every meeting's transcript alongside the rest of your media. Ask answers questions across them, such as “what did we decide about pricing?”, and cites the moments it drew on, each linked to the exact second (how Ask works).

Decisions get one more link. Under each decision in a meeting's brief, Related decisions in other meetings lists up to three similar decisions from earlier meetings and three from later ones, each opening where it was said. We call them related, never replaced, and that wording comes from a measurement. We wrote 18 pairs where a later decision revised an earlier one, and 18 unrelated pairs. At the similarity threshold we ship, MediaFind found 12 of the 18 revisions and linked 2 of the unrelated pairs. No local model we tried could both find revisions and avoid false links. So MediaFind shows the connection and leaves the judgment to you.

When the meeting ends

Stop & save writes the recording and its transcript into your library, where it becomes a meeting like any other; the audio isn't transcribed a second time. Discard throws it all away, and stays free even if a Pro trial runs out mid-meeting, so you're never left with a recording you can't delete. If the app quits in the middle of a meeting, the audio is recovered the next time it starts.

From a meeting's brief, Export .docx and Export .pdf write the notes: overview, action items, decisions and key points, then the transcript with speakers and timestamps. Copy puts the notes on the clipboard as Markdown. Any transcript, meeting or not, also exports as TXT, SRT, VTT, Markdown, Word or PDF. Every file is written on your computer.

What it doesn't do

Meetings is part of MediaFind Pro, and every fresh install comes with a 7-day Pro trial, with no card and no account. Transcript exports and Find & Ask are free.


The best meeting notes are the ones you can check: a transcript you can click, action items in the speaker's own words, and due dates that mean what they meant on the day. MediaFind builds them on your computer, and keeps them there.

Take your next meeting's notes on your own computer

7-day Pro trial. No bot, no account, nothing uploaded.

Download MediaFind MediaFind vs Otter →