A microphone with an audio waveform

Making a Voice Login: The Idea

Why I built VoiceID, a passwordless voice login, and how the frontend records your voice. MIME negotiation, the never-sent prompts, and staged registration.

Sep 27, 2026#voice#nextjs#typescript

Every password I type falls into one of three buckets. Reused, my password manager's 24 characters of nonsense, or forgotten inside a week. I kept wondering what login feels like when the secret is not a string you have to recall, but the way you talk.

This is VoiceID. An account is a username plus a voice. You record five short sentences once, and after that signing in means reading one sentence into your microphone and having it match. There is no password anywhere in it. Not stored, not remembered, not reset over email.

One caveat first, because it shapes everything else. This is speaker verification, not liveness detection. A recording of your voice will pass. The system answers "is this the same voice", not "is a person speaking right now". I have more on that in part three, including the text-dependent challenge I started building and then deleted on purpose.

The idea behind "something you are"

A password has a weakness that never really goes away. The server has to hold proof that you know one. Get that storage wrong and every account that reused that password leaks at the same time.

A voice embedding is not a secret. You cannot hide how you talk, at least not sustainably, so the server is not guarding a secret. It is storing a description of a fact about you.

The trade-offs flip around. People hear you talking all day, so the thing being verified leaks constantly. And you cannot rotate your voice the way you rotate a password. If a voiceprint gets out there is no changing it. This is never going to be the strongest authentication I own. It is convenient in a different way: no typing, nothing to memorise, nothing exposed at a coffee shop. It is also a nice problem to work on, which is reason enough.

The stack

1Browser ── Next.js (:3000)
2                │  rewrite /api/:path*
3                ▼
4            FastAPI (:8000)
5                ├─ ffmpeg          → 16 kHz mono s16le
6                ├─ Silero VAD v5   → keep only real speech
7                ├─ ECAPA-TDNN      → 192-dimensional embedding
8                └─ PostgreSQL + pgvector

The frontend records one blob and sends it. Everything interesting happens behind /api/*, and Next.js proxies that through to the real backend so the session cookie stays first party.

The voice login page: type a username, read the sentence
The voice login page: type a username, read the sentence

Recording audio in the browser

The recorder is a small hook around MediaRecorder. The first thing that annoyed me was that every browser wraps microphone audio in a different container, so the code negotiates a MIME type instead of assuming webm.

useRecorder.ts
1function pickMimeType(): string | undefined {
2  if (typeof MediaRecorder === "undefined") return undefined;
3  const candidates = [
4    "audio/webm;codecs=opus",
5    "audio/webm",
6    "audio/mp4",
7    "audio/ogg;codecs=opus",
8  ];
9  return candidates.find((type) => MediaRecorder.isTypeSupported(type));
10}

Capture constraints stay gentle: mono, echo cancellation on, noise suppression on. I deliberately do not set a sample rate, because browsers ignore that part. Chrome records at 44.1 or 48 kHz whatever you ask for. That is the reason the backend resamples to 16 kHz. Normalising in one place is more reliable than trusting the browser.

useRecorder.ts
1navigator.mediaDevices
2  .getUserMedia({
3    audio: { channelCount: 1, echoCancellation: true, noiseSuppression: true },
4  })

MediaRecorder.start() gets no timeslice, so the whole take arrives as one blob on stop(). Nothing is streamed. Record, stop, send. That is plenty when it is one sentence every now and then.

The client-side "too short" gate

Holding the microphone steady while the page uploads is more annoying than it sounds, so the button drops obviously useless clips before they cost a round trip. A recording under 60% of the minimum length never gets uploaded.

RecordButton.tsx
1recorder.onRecorded((blob, seconds) => {
2  void recorder.cancel();
3  if (seconds < minSeconds * 0.6) {
4    setLastWasShort(true);
5    return;
6  }
7  setLastWasShort(false);
8  void onRecorded(blob, seconds);
9});

The 0.6 is not the real threshold, it is the "that could never have worked" threshold. The server applies the actual rules with actual speech detection, since two seconds of mic time in a loud room can contain almost no voice at all.

Enrolment, one prompt at a time until five samples are recorded

The prompts

There are twelve phrases and every one of them is dull on purpose.

prompts.ts
1export const PROMPTS: readonly string[] = [
2  "The harbour lights were dim that evening.",
3  "We should check the invoice before Friday.",
4  // ...
5  "He explained it twice, and it still made no sense.",
6] as const;

A prompt is there for variety, not security. Five recordings that all say "hi" give you near-duplicate embeddings. Five different sentences give you different phonetics and slightly different pitch, which is something for the matcher to average over. Enrolment walks the prompts in order, so the five samples are always five different phrases. Sign-in picks one at random on each page load.

The part that actually matters: the prompt is never sent to the server. The phrase just gives you something to say. The upload carries the raw audio and nothing else, and the record button says so on it.

The words are only here to give you something to say. They are never sent or checked.

So anyone hoping to map "which phrase is this" onto a canned recording has nothing to map.

One origin for everything

The API runs on a different port than the site, and ports do not share cookies. Browsers split storage by scheme, host and port, so localhost:3000 talking to 127.0.0.1:8000 makes the session cookie a third party. In local development this breaks login in a genuinely confusing way.

The fix is to keep the browser on a single origin. Next.js rewrites /api/... through to the real backend, so every fetch is same origin and the cookie is first party.

next.config.ts
1async rewrites() {
2  return [
3    {
4      source: "/api/:path*",
5      destination: `${API_INTERNAL_URL}/:path*`,
6    },
7  ];
8}

Registration is staged

A username does not exist until a voice does. Registering creates a pending registration: your username plus a random token in a cookie, expiring after an hour. Each recording appends embeddings to that draft. The real user row appears only after the fifth sample, created together with all five voice samples in one transaction.

The result is a structural guarantee: an account with a username but no voice is impossible. The user row and the voice samples either both exist or neither does. Abandon registration halfway and there is nothing to clean up, the draft just expires.

The browser side stops at the blob. Everything past that point is where the work is, so the next two posts are backend. Part two covers the audio pipeline, from untrusted bytes to the cleanest stretch of speech I can pull out of it.