Making a Voice Login: The Backend
Turning an untrusted browser blob into a clean 16 kHz mono speech sample: ffprobe, ffmpeg, Silero VAD, hull trimming, and the failure modes in between.
In the first part we got a blob across the wire. Now the backend has to turn it into a 16 kHz mono sample of pure speech before anything clever can happen. None of this is glamorous, and all of the real difficulty is here.
Assume every upload is garbage
The input is bytes from a browser. It could be webm, mp4, m4a, an onion, a JSON error page that a load balancer spat out, or a text file renamed to .mp3. So the pipeline starts by being extremely cheap about rejecting things, before it spends a second of model time.
The first three gates run on nothing but the raw bytes:
1def _precheck(self, payload: bytes) -> None:
2 self._reject_oversized(payload)
3 if not payload:
4 raise EmptyAudioError("The recording was empty")
5 if not self._looks_like_audio(payload):
6 raise UnsupportedAudioError(
7 "The uploaded file is not a recognised audio container"
8 )_reject_oversized enforces the 10 MB cap. _looks_like_audio is a magic-byte sniff: WAV's RIFF, Ogg's OggS, FLAC's fLaC, MP4's ftyp, EBML for matroska, MP3 sync bytes. Two bytes at the start of a file are enough to tell you it is HTML (<) or JSON ({), and those get rejected immediately. This check used to live in the frontend, until I remembered that nothing in the frontend can be trusted.
Probing before decoding
ffprobe gets asked the cheap questions first: what container is this, what codec, how long is it. Two subprocess invocations before we decode anything.
1def properties(self, payload: bytes) -> AudioProperties:
2 self._precheck(payload)
3
4 command = [
5 self._resolve(self._ffprobe),
6 "-v", "error",
7 "-print_format", "json",
8 "-show_format", "-show_streams",
9 "-i", "pipe:0",
10 ]
11 completed = self._execute(command, payload, "audio inspection")The properties get checked against two allowlists: supported containers (webm, matroska, mp4, mov, ogg, wav, mp3, ...) and supported audio codecs (opus, vorbis, aac, mp3, flac, pcm_s*, ...). If any header says the file has no audio track, or the codec is one we do not handle, it is rejected with a 422. That one is the client's fault, not ours.
Probing first is what makes the duration gate cheap. If someone recorded a 40-minute video and helpfully attached audio to it, ffprobe tells us before ffmpeg spends time decoding it.
The cut: 16 kHz mono, nothing else
The decode command is where the browser blob becomes the thing the model wants.
1command = [
2 self._resolve(self._ffmpeg),
3 "-hide_banner", "-nostdin", "-loglevel", "error",
4 "-i", "pipe:0",
5 "-vn", "-sn", "-dn",
6 "-ac", str(TARGET_CHANNELS),
7 "-ar", str(TARGET_SAMPLE_RATE),
8 "-c:a", "pcm_s16le",
9 "-f", "s16le",
10 "pipe:1",
11]Four things happen at once. Video streams get dropped, audio is downmixed to one channel, any source sample rate is resampled to 16 kHz, and the output is raw signed 16-bit PCM. No intermediate file, no temp directory: ffmpeg reads pipe:0 and writes pipe:1. Nothing touches the filesystem, which also means nothing is left on disk after a failed attempt.
Everything downstream can then assume int16, 1-D, 16000 Hz. That invariant is enforced at construction time:
1def __post_init__(self) -> None:
2 if self.samples.dtype != SAMPLE_DTYPE:
3 raise ValueError("NormalizedAudio expects int16 samples")
4 if self.samples.ndim != 1:
5 raise ValueError("NormalizedAudio expects mono (one dimensional) samples")Why 16 kHz? It is an accident of the model. ECAPA-TDNN from SpeechBrain is trained on 16 kHz audio, and I picked a pretrained model instead of training one. The frontend cannot be trusted to deliver that rate, and you cannot reliably ask browsers for it anyway, so the pipeline owns it. Every container, codec and sample rate in the world funnels into one canonical format at a single chokepoint, which saves a lot of per-browser special-casing.
The silence check is dumb but correct:
1@property
2def is_silent(self) -> bool:
3 if self.samples.size == 0:
4 return True
5 return not bool(np.any(self.samples))If every sample is literally zero it is silence. A muted microphone produces a stream of zeros. We check this after decode even though VAD would catch it anyway, because it is free and it handles the "I forgot to unmute" case.
Actually finding the voice: Silero VAD
The audio is 16 kHz mono int16 now, but it is still a whole recording, probably with room tone before you started talking and a tail after you stopped. The voice activity detector carves out the speech.
The VAD is Silero v5, loaded as ONNX:
1self._model = load_silero_vad(onnx=True, sampling_rate=SAMPLE_RATE)I went with Silero instead of the usual webrtcvad for two reasons. It is a neural network that scores every 32 ms frame with a probability of speech rather than comparing against an energy threshold, and it is much better in noisy rooms. It also runs through ONNX, so it is the one model in the stack that does not want the torch backend loaded at all.
The segmenter runs with Silero's standard knobs: 200 ms minimum speech and silence, plus a 50 ms pad on either side of a speech run so the first consonant does not get clipped.
get_speech_timestamps has one wrinkle: it is stateful. The model keeps recurrent state between calls, so calls have to be serialized or you get garbage. The detector holds a threading lock around every VAD call:
1with self._lock:
2 timestamps = get_speech_timestamps(
3 tensor, model,
4 sampling_rate=SAMPLE_RATE,
5 threshold=SPEECH_PROBABILITY_THRESHOLD,
6 min_speech_duration_ms=MIN_SPEECH_DURATION_MS,
7 min_silence_duration_ms=MIN_SILENCE_DURATION_MS,
8 speech_pad_ms=SPEECH_PAD_MS,
9 return_seconds=False,
10 )Trimming, the "hull" version
The instinct is to take every segment the VAD returned and concatenate them. It is the wrong instinct. Cutting between two syllables that sit 250 ms apart turns one word into two clipped stumps. ECAPA-TDNN is spectral and trained on whole utterances, and it copes with a little silence inside its input far better than with a hard cut through a phoneme.
So trimming keeps everything from the start of the first segment to the end of the last one:
1def trim_to_speech(samples: np.ndarray, result: VadResult) -> np.ndarray:
2 if not result.segments:
3 return np.empty(0, dtype=samples.dtype)
4 start = max(0, result.segments[0].start_sample)
5 end = min(samples.size, result.segments[-1].end_sample)
6 if end <= start:
7 return np.empty(0, dtype=samples.dtype)
8 return samples[start:end]The head and tail silence goes, pauses in the middle stay. It is a deliberate trade: a sliver of noise between the words is worth it so that no word gets chopped.
Minimum viable voice
A speaker embedding is only meaningful if there is enough speech to embed. Whisper for 0.2 seconds and the vector you get out is mostly room tone and microphone noise. So there are two floors:
- 0.8 seconds of total speech across the whole clip
- a longest single segment of at least 0.7 seconds
The error messages carry the real numbers, because someone who hears "that sentence was not long enough" gets annoyed in a way that is not useful to debug, while someone who reads "only 0.5s of speech was detected, which is below the 0.8s minimum" can just try again:
1def require_sufficient_speech(result: VadResult, settings: Settings) -> None:
2 if not result.has_speech:
3 raise InsufficientSpeechError("No speech was detected in the recording")
4 if result.speech_seconds < settings.min_total_speech_seconds:
5 raise InsufficientSpeechError(
6 f"Only {result.speech_seconds:.1f}s of speech was detected, "
7 f"which is below the {settings.min_total_speech_seconds:.1f}s minimum"
8 )
9 if result.longest_segment_seconds < settings.min_speech_seconds:
10 raise InsufficientSpeechError(
11 f"The longest stretch of speech was {result.longest_segment_seconds:.1f}s, "
12 f"which is below the {settings.min_speech_seconds:.1f}s minimum"
13 )The full order
Everything hangs together in AudioPipeline.process_audio, which is deliberately a straight line:
- reject empty or oversized
- ffprobe, reject too long
- ffmpeg decode to 16 kHz mono int16
- reject silent
- VAD
- enforce minimum speech
- trim to the outer hull
- hand back a
ProcessedAudiowith the trimmed samples
Every failure in there maps to a 422 with a readable message, because every one of them is the client's fault. When the infrastructure fails instead, a model that will not load or a database that will not connect, the error becomes a 503. Keeping those two apart is what lets the API say "your recording was silent" instead of "internal server error" every time the GPU box hiccups.
What comes out the other end is a short stretch of voice with the mic hum mostly gone, which is exactly what the embedding model wants. That is the whole flow working end to end:
Part three turns that stretch of speech into a number that decides whether you get in.