Comparison
Meeting notes without a bot: how to record a meeting in the same room
Search for an "AI meeting assistant" and almost every tool that comes back does the same thing: a bot joins your Zoom, Meet or Teams call. So what happens when the meeting is around a table and there is no link to join?
The category quietly makes an assumption
Fireflies, tl;dv, Fathom, Read, Otter — every one of them starts its flow with a calendar invitation. The bot joins the link, records, produces a transcript. That is good design: the bot receives each participant's audio on a separate channel, so it never has to guess who is speaking. The channel already says who.
But a large share of meetings in Turkey still happen around a table: a client visit, a supplier negotiation, an interview, a class, a consulting session, a meeting with a lawyer. These meetings have no link. None of the category-leading tools is any use in that room.
Why recording one room is technically harder
The difference in one sentence: the bot has one audio channel per person, the phone on the table has a single mixed channel.
Three difficulties follow:
- The speaker has to be found from the sound. There is no channel; you have to derive who is speaking from the audio itself. That work is called diarization, and it is the genuinely hard part.
- Overlapping speech. In the same room people cut across one another. In an online meeting that is rarer, and even when it happens the channels are separate.
- Distance and room acoustics. A headset microphone sits next to your mouth; the phone on the table hears you from two metres away, with the echo attached.
If a tool says "I do speaker separation", the question is under what conditions. A tool that reads the channel in an online meeting will not give you the same result on your recording at the table.
The real options for a face-to-face meeting
1. Record audio, paste it into an AI afterwards
Record with the phone's Voice Memos, hand the file to a transcription service, then hand the text to ChatGPT. It is free and it works — up to a point.
Where it jams: there is no speaker separation, so "who promised what" is lost. A two-hour recording runs into upload size limits. And most importantly, you have to remember to do those three steps after the meeting. That is why most recordings sit on the phone, never transcribed.
2. Tools with an on-device recording mode
Some tools, like Notta and Transkriptor, have a path other than the bot: recording directly from the browser or the mobile app. These can be used for a face-to-face meeting, and both support Turkish.
What to expect: after the recording ends the file goes to a server, is transcribed there, and the speaker separation is done there too, specific to that file. There is no screen being transcribed live and no speaker identity carried from one recording to the next. Plans are sold by the minute.
3. An app written for the meeting at the table
Ses Notu is that third path. You put the phone in the middle of the table and press record:
- Speech turns into text on screen while people are talking — not when the recording ends.
- Speakers are separated: Speaker 1, 2, 3. Names said during the meeting are matched to the right voice. The voice prints stay on the phone, and the same person is recognised in later meetings.
- When the recording ends, the summary, the decisions taken and the action items with an owner and a deadline are ready. "By Friday" is turned into a real date and goes to Reminders with one tap.
- Recording continues while the screen is locked and resumes where it left off after an incoming call. You can start and stop it from the Apple Watch.
Speaker separation runs inside the phone: the CAM++ voice print model (WeSpeaker, Apache-2.0) on ONNX Runtime, a ~28 MB file shipped inside the app. A voice print is biometric data and never leaves the phone. How it was done.
| Approach | Works in one room | Live | Speaker separation | Output |
|---|---|---|---|---|
| Bot-based tools Fireflies, tl;dv, Fathom, Read | No — needs a link | Yes | From the channel, exact | Meeting note |
| Voice Memos + ChatGPT | Yes | No | None | Raw text |
| On-device recording mode Notta, Transkriptor | Yes | No | On the server, per file | Transcript plus summary |
| Ses Notu | Yes — that is the design | Yes | On the device, persistent voice print | Summary, decisions, action items |
Before you record: tell the room
Bot-based tools have a side benefit — a participant appears on screen saying "recording", and everyone sees it. When you put a phone on the table there is no such visual warning.
In Turkey, under KVKK (the Turkish data protection law), an audio recording counts as personal data and a voice print as special-category personal data. Telling the room before you start the recording is required by the legal side and by the relationship alike. Ses Notu reminds you of this inside the app too, but the responsibility belongs to whoever is recording.
Ses Notu — takes the notes for the meeting at the table. No bot, no Zoom. iPhone and Apple Watch. ₺199.99/mo, first week free. Download on the App Store