All posts

Live captions: exactly where your meeting's audio goes

In a Nemi meeting, your camera and your voice go from your browser to the other people in the call, and there is no media server of ours that can see or hear them. Live captions are the one feature that sends what is said to a company outside the call, and an exception is only acceptable if it is exact. This is where the audio goes when the host presses the button, where the words go after that, and how a voice number becomes a name without anybody's voice being identified.

Privacy · · 10 min read

In a Nemi meeting, your camera and your voice go from your browser to the other people in the call. There is no media server of ours that can see or hear them. Live captions are the one feature that sends what is said to a company outside the call, and when a product makes an exception to its own rule the least it owes you is the exact route.

Here it is: when the host turns captions on, the host's browser mixes everybody's microphone into one stream and sends it to Soniox, a speech recognition company, in its European Union region. That audio never passes through a Nemi server. It does leave the host's browser, and it is heard by a company that is not in the meeting. Everything below is the detail of that one sentence.

How do live captions work in a video call?

Every live caption system does the same three things: it streams the audio to a speech recogniser, gets text back while people are still talking, and shows it. What differs is where the audio is sent from and who receives it. Start with the call itself. Each pair of people has its own encrypted connection, browser to browser. On networks that refuse direct connections, usually corporate ones, the encrypted stream goes through a relay we run, which forwards packets it cannot decrypt and keeps none of them. That is true with captions on or off. Then switch captions on and watch what is added.

A four person call, with and without captions

audio in, words outSonioxspeech service, EU regioncaption text, passed on and not storedNemithe meeting's chat channelAAnnahostDDev (guest)CChloeBBen
The call itself
Browser to browser, encrypted between each pair. No server of ours is on these lines.
Caption audio
One stream, from Anna's browser straight to Soniox in the EU. It does not pass through Nemi, and it does leave Anna's browser.
Caption text
Anna's browser posts each line to the room. It travels the same way the chat does, and we pass it on without storing it.
Anna hosts, so the captions switch is hers. Dev joined from a link without an account; his voice is in the mix like everybody else's. Toggle Chloe's network to see where the relay sits.

Three new lines appear, and they carry different things. The thick one is audio: one stream, sampled at 16 kHz and compressed with Opus, sent in quarter second pieces so a word can be on screen while the sentence is still being spoken. It goes from Anna's browser to Soniox's EU endpoint and to nowhere else. What comes back is text. Anna's browser then posts each line to the room over the same channel the meeting chat uses, and our server passes it to everybody else without writing it down.

So the honest summary has two halves. The audio never touches our servers, and the words do pass through them, briefly, on their way to the other people in the call. We keep neither the audio nor those lines; the only lasting copy of the words is the transcript in the host's meeting notes. The privacy policy says both in section 3.9, and names Soniox in its list of the companies we use. Nemi AI makes a similar trade for files, and its post sets out where those go.

Why does the host decide for everybody?

Because one stream is the only shape in which this works well. An earlier design had each person caption their own microphone in their own browser. It sent nothing anywhere, and it could not do the job: many browsers have no recogniser of their own, and three people captioning produced three partial records of one conversation that could not be put together.

One mixed stream is one record, one bill and one switch. It is also the only way the service can tell voices apart, because two voices have to be in the same audio to be compared at all. The cost of that is real: when the host switches it on, everybody's words go to the service, including a guest's, including whatever was said before somebody noticed. That is why only the host has the switch, and why every participant sees a Captions on marker from the moment it is pressed, not from the first word. If you would rather not be transcribed, the moment to say so is in the meeting.

What happens when the host presses Turn on captions

Host's browser to Nemistart captionsNemi to Soniox, EUtemporary keyHost's browser to Soniox, EUmixed audioSoniox, EU to Host's browserwordsHost's browser to Everybody elsenamed lines, via NemiHost's browser to Nemistill on (every 30 s)Host's browsertranscript into the host's notes

1/5Nemi checks the host's hours and asks Soniox for a temporary key

Our own key to the service never reaches a browser. The host gets a key that must be used within minutes, for a stream that Soniox itself will close after an hour, or sooner if the host has less of the month's allowance left.

How does a transcript know who said what?

The speech service separates voices, but it only knows them as numbers: voice 1, voice 2. A transcript that says “Speaker 2 will send the list on Thursday” is not much use next week. The meeting, meanwhile, already knows who is talking at any moment, because the host's browser measures every participant's audio level to decide whose tile to highlight.

So the two are lined up on the same clock. Each finished line asks who the level meter said was speaking during it, and adds its length to a tally for that voice. The voice's name is whoever leads the tally. One line can be matched wrongly, when people talk over each other or the meter hands over a beat late. The total over a meeting is much harder to get wrong, and when it changes its mind, lines already written are corrected.

Matching voices to names, line by line

Captions, as they appeared

  1. Anna: Right, the launch. Three weeks, and the press list is not done.
  2. Anna: I can take the press list.

The tally, per voice

  • Voice 1Anna

    Anna 8.3 s

  • Voice 2Anna

    Anna 3.6 s

The meeting notes, written at the end

Anna: Right, the launch. Three weeks, and the press list is not done. I can take the press list.

A minute of an invented stand up, run through the same matching code the host's browser runs. Scrub to 13 seconds and then past 24 to see the second voice change its name.

Two cases in that replay are worth stopping on. Ben starts talking while the meter still has Anna as the speaker, so his first line appears under her name. His next, longer line outweighs it, the tally flips, and the notes say Ben for both. Then there is the quiet question at 33 seconds, asked while the meter had nobody: there is no evidence for any name, so it stays Speaker 4. That is a poor label and an honest one. Guessing a participant's name would put words in somebody's mouth, in a document the host keeps.

An unmatched voice stays Speaker 4. A guess would put words in somebody's mouth.

None of this identifies a voice. There is no voiceprint, nothing is compared across meetings, and the tally is a sum of milliseconds that lives in the host's tab and is gone when it closes. What leaves that tab is the name it arrived at, on each line.

How many hours of captions do I get?

Captions cost money per minute the moment they run, so they are counted, in hours per calendar month, against the account that hosts the meeting. Joining somebody else's meeting never uses your hours, and a guest has no account to count anything against. The clock starts when the audio does, so the seconds spent asking for a key are not charged.

Hours of live captions per month, by plan

Hours per month

Free

0

Starter

5

Creator

10

Pro

10

Max

10

Business

100
Read from the plan definitions the server enforces. The full comparison is on the pricing page.

When the hours run out, captions stop and nothing else does: the meeting carries on, and the message names whose plan the hours belonged to. The allowance comes back at the start of the next month. See pricing for the plans, and the help page on captions for where the switch is.

What is kept, and by whom?

By us: a count of seconds of captioning per account for the current month, which is a number and nothing else, and the meeting notes, which are the host's own document. We do not hold the audio at any point, and we do not store the caption lines that pass through on their way to the room.

By Soniox: what its own terms allow for audio it is sent, inside its EU region, where both the processing and any storage stay. That part is governed by its terms, not by our retention table, and we say so rather than pretend otherwise.

Questions people ask

How do live captions work in a video call?

The call's audio is streamed to a speech recognition service, which sends back text while the sentence is still being spoken, and that text is shown as captions. Where the audio is sent from, and to whom, differs between products. In Nemi Meet the host's browser mixes everybody's microphone into one stream and sends it to Soniox in the EU, and the words come back to the host's browser.

Is my voice sent to a third party when captions are on?

Yes. When the host of a Nemi meeting turns captions on, everybody's audio, guests included, is sent from the host's browser to Soniox, a speech recognition company, in its European Union region. It does not pass through Nemi's servers, but it does leave the host's browser. Every participant sees a Captions on marker for as long as it runs.

Who can turn on live captions in Nemi Meet?

Only the host of the meeting. One mixed stream is one record, one bill and the only way voices can be told apart, so the switch belongs to one person. If you would rather not be transcribed, the moment to say so is in the meeting, while the Captions on marker is showing.

Are live captions saved after the meeting?

When the host leaves, the finished sentences are written into the meeting notes, an ordinary document in the host's workspace that outlives the meeting. Nemi does not store the caption lines that pass through its servers on the way to the other participants, and never holds the audio. Soniox handles the audio under its own terms, inside its EU region.

How many hours of live captions does each Nemi plan include?

Free includes none, Starter 5 hours a month, Creator, Pro and Max 10 hours, and Business 100 hours. The hours belong to the account that hosts the meeting, so joining somebody else's meeting never uses yours. When they run out, captions stop and the meeting carries on.

How does a meeting transcript know who said what?

The speech service separates voices but only knows them as numbers. Nemi lines those numbers up against who the meeting's audio levels said was speaking at the time, in the host's browser, and gives each voice the name it matched for the longest time. A voice that matches nobody stays Speaker 2 or Speaker 4 rather than getting a guessed name.

Where the numbers come from

  • Mixing and streaming the audio from the host's browser: lib/meet/transcribe.ts
  • The EU endpoints and the temporary key: lib/meet/soniox.ts (SONIOX_WEBSOCKET_URL, createTemporaryKey)
  • Turning voices into names, and the transcript: lib/meet/transcription.ts (nameDuring, SpeakerNames, speakerLabel, buildTranscript)
  • How the matching is called during a meeting: components/meet/useMeetRoom.ts
  • The 30 second tick and the monthly count: lib/meet/plan.ts (CAPTION_TICK_INTERVAL_MS), lib/meet/captions-usage.ts
  • Caption hours per plan: lib/pricing-plans.ts (meetCaptionHoursPerMonth)
  • What is sent, to whom, and what is kept: /privacy, sections 3.9 and 6.1

Read next

Use the thing we write about.

Files, docs, sheets, photos, calendar and meetings in one account.