Guide
Who said what in a Turkish meeting: how speaker separation works
The phrase "it separates speakers" is used for three different things, and the three differ enormously in difficulty. Knowing which one you need in a Turkish meeting is half of choosing a tool.
Three separate jobs, one name
- Diarization. Working out which part of the audio was said by whom. The output is labels of the form "speaker 1 / speaker 2"; there are no names. It is language-independent — it looks at the voice print, not the word. Turkish does not make this step harder.
- Identity. Knowing that "speaker 2" in this recording is the same person as "speaker 1" in last week's. That requires storing a voice print unique to the person. Also language-independent, but it means biometric data.
- Name assignment. Working out that "speaker 2" is Ayşe. This one is language-dependent, because the only source for the name is the conversation itself: "Ayşe Hanım, what do you think?"
Most tools in the category solve the third step from the participant list: a bot joins the online meeting, everyone's audio arrives on a separate channel, and the name is already written in the participant list. In a meeting in one room there is no such list — one microphone, mixed audio, no names.
What makes name assignment hard in Turkish
The way you address someone by name in Turkish does not resemble English. "Ayşe Hanım", "Ahmet Bey", "as Mehmet Bey said", "hocam" (teacher), "müdürüm" (boss) — some of these carry a name, some do not, and some are about a third person entirely. Then there are the suffixes: "Ayşe'nin", "Ayşe'ye", "Ayşeciğim" all point at the same person, but plain string matching cannot catch them.
That is why trying to solve name assignment with pattern matching is fragile in Turkish. You have to read the meeting as a whole and make the inference "this line addresses Ayşe, therefore the person answering is Ayşe" — which is the work of a model that sees context.
One caveat: a name can only be found if it is said in the conversation. In a meeting where nobody addresses anyone by name, no tool can know the names. That is not a quality problem; the information is not in the recording.
The real limits of a single microphone
When you put the phone on the table and record, three things make diarization harder, and it is worth knowing them in advance:
- Overlapping speech. When two people talk at once, single-microphone diarization does not do well. The moments where a meeting turns into an argument are the moments where the transcript is weakest.
- Distance. The person at the far end of the table arrives both quieter and with more room echo. The same person's voice can produce a different voice print close to the microphone and far from it.
- Similar voices. Two people of the same gender with a similar pitch and tempo are the classic case diarization confuses.
The fix is simple and not technical: put the phone in the middle of the table, screen up, with no paper or bag on top of it.
How Ses Notu does it
The three steps run in three different places:
- Diarization happens live while you talk; the transcript grows on screen line by line, already split by speaker.
- Identity is inside the phone. The model that extracts the voice print (CAM++, ONNX Runtime, ~28 MB) ships inside the app; the vectors it produces never leave the device and can be deleted from settings. Name a person once and they are recognised automatically in later meetings. How it works — and a bug along the way that gave wrong answers without ever raising an error — is described in this post.
- Name assignment happens when the meeting ends, by looking at the transcript as a whole: if names came up in the conversation, it infers which speaker is who and relabels accordingly. It does not always get it right, which is why you can rename any speaker by hand in one tap, and the correction is applied across the whole transcript.
And the missing side: Ses Notu only runs on iPhone and Apple Watch, and it does not put a bot into online meetings. If your meetings happen in Zoom or Teams, read the distinction on this page; there are tools better suited to that job.
Ses Notu — takes the notes for the meeting at the table and separates the speakers. iPhone and Apple Watch, Turkish and English. ₺199.99/mo, first week free. Download on the App Store