# How does Canto tell speakers apart?

> Canto separates speakers by voice into labelled turns, then AI infers real names from the conversation. You can rename any speaker manually.

Updated: 2026-08-16 · Canonical: https://usecanto.com/help/transcription/speaker-separation/

During transcription, Canto separates the audio by voice, so each transcript line belongs to one speaker's turn. Speakers start out labelled **Speaker 1**, **Speaker 2**, and so on; after the recording is indexed, AI tries to replace those with real names it heard in the conversation.

## How are speakers separated?

Speaker separation (diarization) happens automatically during transcription — there is no setting and no fixed speaker count to declare. The transcript is built as a sequence of turns: each time a different voice takes over, a new line starts under that speaker's label, with its own timestamp. Long monologues are split into shorter lines for readability, still under the same speaker.

Separation is based purely on how voices sound in the recording. It works best when people speak one at a time into a reasonably close microphone; heavy crosstalk, similar-sounding voices, or a distant phone on a big conference table make it harder. See [Get the best audio quality](https://usecanto.com/help/recording/microphone-tips/).

## Where do speaker names come from?

Three sources, in order of priority:

1. **Names you set** — a manual rename always wins. See [Rename speakers in a transcript](https://usecanto.com/help/transcription/rename-speakers/).
2. **AI-inferred names** — while preparing the recording (the **Preparing AI** stage), Canto reads the conversation for introductions and direct address ("Thanks, Maria — over to you") and assigns real names where it's confident. This can take a minute after the transcript first appears, so labels may upgrade from **Speaker 2** to a name shortly after.
3. **The fallback** — **Speaker 1**, **Speaker 2**, … in order of first appearance.

## Do notetaker meetings get real participant names?

Meetings recorded by the Canto notetaker bot have the participant list from the call itself, so speakers are matched to the actual meeting participants rather than inferred from what was said. This is the most reliable naming path — see [What is the Canto notetaker?](https://usecanto.com/help/notetaker/what-is-the-notetaker/)

## Can I fix a wrong label?

Yes. Renaming a speaker applies to every line from that voice, for everyone with access to the recording, and you can revert to the automatic name at any time — see [Rename speakers in a transcript](https://usecanto.com/help/transcription/rename-speakers/).

What renaming can't fix is a separation error: if two people were merged into one speaker, or one person was split across two labels, that happened at transcription time and there is no way to reassign individual lines to a different speaker. See [Speakers are mixed up or mislabeled](https://usecanto.com/help/troubleshooting/speakers-mislabeled/) for what helps on the next recording.

## How many speakers can Canto handle?

There is no hard cap you need to think about — normal meetings, interviews, and group conversations are fine. Accuracy degrades gracefully as a conversation gets more crowded and overlapping rather than stopping at a fixed number.

Related articles:

- [Rename speakers in a transcript](https://usecanto.com/help/transcription/rename-speakers/)
- [What is the Canto notetaker bot?](https://usecanto.com/help/notetaker/what-is-the-notetaker/)
- [Speakers are mixed up or mislabeled](https://usecanto.com/help/troubleshooting/speakers-mislabeled/)
- [How do I get the best audio quality?](https://usecanto.com/help/recording/microphone-tips/)
