Making a Voice Login: The Backend, Part 2
ECAPA-TDNN embeddings, cosine similarity, the median aggregation, the hand-picked threshold, and how sessions and rate limits hold the door.
In the previous part we squeezed a browser blob down to a clean 16 kHz mono sample of speech. This is the actual recognition: turning that waveform into a number, comparing it against what we stored at enrolment, and deciding whether it is the same voice.
The model: ECAPA-TDNN
The whole system leans on one pretrained model, SpeechBrain's spkrec-ecapa-voxceleb, trained on VoxCeleb. It is an ECAPA-TDNN: a convolutional network that reads mel spectrograms and squeezes a whole utterance into a single vector called an embedding.
The front end is a standard SpeechBrain feature extractor, 80 mel filters over a 25 ms window with a 10 ms hop, on 16 kHz audio. That is the format everything else was normalized to in the previous part. Then a stack of three SE-Res2Net blocks with increasing dilation, attentive statistics pooling, and a final layer projecting down to 192 dimensions.
Those 192 numbers are the whole description of what a voice sounds like. The same speaker reading a different sentence gives very similar embeddings. Different speakers land far apart in that 192-dimensional space. Everything after this point is geometry.
Loading the model, reluctantly
ECAPA is about 85 MB of checkpoints sitting in the Hugging Face hub. We do not want that loaded at import time, on every worker, or worse on every request. So the speaker service loads lazily, behind a double-checked lock, so only one thread does the work even if five requests land at once:
1def _ensure_model(self) -> EncoderClassifier:
2 if self._model is not None:
3 return self._model
4 with self._lock:
5 if self._model is not None:
6 return self._model
7 try:
8 from speechbrain.inference.speaker import EncoderClassifier
9 model = EncoderClassifier.from_hparams(
10 source=self._settings.voice_model,
11 savedir=self._settings.voice_model_cache_dir,
12 run_opts={"device": device},
13 )
14 model.mods.eval()
15 except Exception as exc:
16 # ... wrapped as ModelUnavailableError so the API answers 503The first request pays the download and load cost. In the Docker deployment that first wait is long enough to matter, which is why the container health check allows a 15-minute start period. A cold start downloads roughly 85 MB.
The odd part: probing the embedding dimension
The 192 is not hardcoded anywhere. The first time the model loads, we ask it how wide its output is by feeding one second of silence through it:
1PROBE_SAMPLES = 16_000
2
3def _probe_dimension(self) -> int:
4 with torch.no_grad():
5 probe = torch.zeros(1, PROBE_SAMPLES)
6 lengths = torch.ones(1)
7 output = model.encode_batch(probe, wav_lens=lengths, normalize=False)
8 return int(output.shape[-1])Why not just write 192? Because the number belongs to the checkpoint, not to my code. If someone swaps in a different model, a 256-dim x-vector say, the dimension should follow on its own. Probing turns the embedding width into a runtime fact rather than a magic constant that silently disagrees with the database column. It goes further than that: the migration sizes the vector column from the same probed value, so the schema cannot drift away from the loaded model.
From waveform to embedding
The embedding itself is three dull steps. int16 PCM becomes float, scaled by 1/32768 to land roughly in [-1, 1]. That goes through the model as a batch of one. The output vector gets L2-normalized to unit length.
1def generate_embedding(self, samples: np.ndarray) -> Embedding:
2 if samples.size == 0:
3 raise ModelUnavailableError("Cannot embed an empty audio buffer")
4 model = self._ensure_model()
5 waveform = torch.from_numpy(np.ascontiguousarray(samples.astype(np.float32)) / 32768.0)
6 batch = waveform.unsqueeze(0)
7 lengths = torch.ones(1)
8 with torch.no_grad():
9 output = model.encode_batch(batch, wav_lens=lengths, normalize=False)
10 vector = output.squeeze().cpu().numpy().astype(np.float32)
11 norm = float(np.linalg.norm(vector))
12 if norm > 0.0:
13 vector = vector / norm
14 return [float(value) for value in vector.tolist()]Notice normalize=False. That is deliberate. SpeechBrain wants to apply a global normalization it learned on VoxCeleb before comparing embeddings. I skip it and do my own L2 normalization instead, because the comparison never uses SpeechBrain's cosine implementation. I compute the similarity myself, so I want the vectors shaped the same way I compare them.
The L2 step matters more than it looks. It takes volume out of the picture: a whispered voice and a shouted one produce colinear embedding directions, so once normalized they land in the same place. How loud someone was should not change who they are.
Comparing: cosine
With two unit vectors, "how close" reduces to the cosine of the angle between them. A single float in [-1, 1], computed with nothing more than a dot product:
1def _cosine(left: np.ndarray, right: np.ndarray) -> float:
2 denominator = float(np.linalg.norm(left) * np.linalg.norm(right))
3 if denominator == 0.0:
4 return 0.0
5 return float(np.dot(left, right) / denominator)Two embeddings from the same voice score high. Two different voices score lower. What is left is how to combine several scores and where to draw the line.
Aggregation: why five samples, and why median
Enrolment collects five samples. Not because five sounds good, but so the comparison has a chance of being stable. One recording of your voice is noisy. Five readings cover drafts, colds, angry Monday mornings.
The comparison computes a cosine against every enrolled sample, then combines them. The default strategy is the median:
1def _median(scores: Sequence[float]) -> float:
2 return float(statistics.median(scores))With five samples the median is literally the third-best score of five. That one choice makes the system tolerant of a single bad enrolment. If one of your five samples was recorded at a train station, the median ignores it. I left four other strategies available behind a setting, mean, trimmed mean, minimum and centroid, but median is the default because it needs no tuning and survives one bad apple.
The threshold I refuse to pretend is scientific
The decision is one line:
1def decide_authentication(*, speaker_score: float | None, threshold: float) -> bool:
2 if speaker_score is None:
3 return False
4 return speaker_score >= thresholdAnd the threshold is 0.35, picked with my gut rather than a graph. The honest way to calibrate is to run a batch of genuine and imposter examples through the comparison, plot the error rates, and pick the point where false accepts and false rejects cross. I have not done that, because building a labelled dataset for a hobby project is a real chunk of work. SpeechBrain's own default is 0.25. I set mine stricter because the enrolment data here is short and captured on consumer microphones, so borderline matches deserve suspicion.
What I did instead was log every score. Each attempt, pass or fail, username real or not, writes its speaker_score to the database. Once I have a real evaluation set those scores become the threshold, and the plumbing is already sitting in the login_attempts table.
The login route, in order
The shape of the login endpoint is more careful than it looks:
1limiter.check_attempt(get_client_key(request)) # rate limit first
2normalized = normalize_username(username)
3# lockout gate, checked BEFORE looking up the user or decoding audio
4since = now - dt.timedelta(seconds=settings.lockout_window_seconds) # 300s
5if attempts.count_recent_failures_for_username(...) >= settings.max_voice_failures: # 10
6 raise AuthenticationLockedOutError("Too many failed voice attempts. Try again later.")
7
8user = users.find_by_username(normalized)
9if user is None:
10 attempts.record(accepted=False, speaker_score=None) # log it anyway
11 raise InvalidCredentialsError(...) # then rejectRate limit first. Then the lockout gate, before we know the user exists and before we spend a second decoding audio. Then if the username does not exist we still record a failed attempt and return exactly the same 401 shape as a voice mismatch. Very little gives itself away. An attacker cannot tell "no such user" from "wrong voice" by looking at the response. The rule I settled on is to skip expensive work, like furiously decoding an uploaded file, when a cheaper check has already failed. It keeps load down under attack.
A voice mismatch returns a deliberately bland message:
1VoiceVerificationFailedError(
2 status_code=401,
3 code="voice_verification_failed",
4 message="Voice verification failed. Please try again.",
5)No score, no threshold, no hint about which sample matched. Everyone gets the same sentence.

What this is not, and what I deleted
None of this makes VoiceID safe against a recording. I said so in part one and I built the evidence the hard way.
The first migration created voice_challenges and voice_attempts tables, a text-dependent challenge system where each login would get a fresh phrase and the backend would score both what you said and who said it. Migration 0002 drops those tables. I pulled it out on purpose. It doubled the surface area, it needed a full speech-to-text model to score the phrase, and, the actual reason, I was not building a surveillance system. A login that accepts your recorded voice is a limitation worth documenting, not a problem to paper over with theatre. The README says it plainly: replaying a recording will pass, this is verification and not a liveness check.
In the same later migration passwords got removed entirely. The backend never stores anything resembling a secret. No raw audio, no transcripts, no password hashes. What is stored is exactly:
voice_samples, 192 floats per sample, per userlogin_attempts, the scores, for calibration and abuse monitoringsessions, a token whose SHA-256 hash is stored rather than the token itself
Tokens and cookies
The session token is 32 bytes from secrets.token_urlsafe, handed to the user once as a cookie, and stored only as a SHA-256 hash. There is no bcrypt, no pepper, no key rotation. The code comment explains why: bcrypt exists to slow down brute-force guessable passwords, and a 256-bit random token does not need slowing down. The hash is there so that if the sessions table leaks, the token still cannot be replayed.
The cookie is HttpOnly so JavaScript never sees it, and SameSite=lax so browsers will not send it on cross-site POST or DELETE. In production the app refuses to start unless SESSION_COOKIE_SECURE is true:
1if self.environment == "production" and not self.session_cookie_secure:
2 raise ConfigurationError(
3 "SESSION_COOKIE_SECURE must be true in production, otherwise the "
4 "session token is sent in clear text over the network"
5 )A config error at boot is better than a security hole in production.
The doorbell: rate limits and lockout
Two sliding windows cover the login door:
- Rate limit: 30 login attempts per client in 300 seconds, on top of a 120-requests-per-60-seconds cap for the whole API.
- Lockout: any single username that fails 10 voice verifications inside 300 seconds is frozen with "Try again later."
The rate limit is a deque of timestamps per client key, kept in process memory. That is fine for one box, since the container runs a single worker, and it is a real limitation if you scale horizontally, because state living in one process stops being shared state at two. The Retry-After value is even computed and then never sent as a header, which I keep meaning to fix. The client key is the first X-Forwarded-For hop only when the proxy is trusted, otherwise it falls back to the socket address.
Scoping lockout per username rather than per client has a side effect I accepted on purpose. An attacker can lock you out by failing as you deliberately. It is a denial-of-service option, but the alternative, no per-username lockout, lets them try ten thousand variations of your voice instead. Being wrong there is worse.
Where it ends
The system is honest about its ceiling. It is voice verification: the same voice across enrol and test, with a score, a threshold and a sliding window in front of it. It is not authentication for anything that deserves a liveness check, and nobody should bank on this. Given that, naming the limitation seemed better than burying it, logging every score for later, and keeping the pipeline dull enough that swapping the model or calibrating the threshold is a config change rather than a rewrite.
The model, the pipeline and the whole plan are in the GitHub repo, and there is a summary of the project itself on the work page.