Skip to content

feat(voice-detect): report the encoder family so audio-registered voices are fingerprinted - #12499

Merged
mudler merged 3 commits into
masterfrom
feat/voice-detect-encoder-family
Oct 8, 2026
Merged

mudler merged 3 commits into
masterfrom
feat/voice-detect-encoder-family

Conversation

@mudler-agent

Copy link
Copy Markdown
Collaborator

Description

Voices registered from audio through the voice-detect backend (POST /v1/voice/register with audio) had no encoder fingerprint, because libvoicedetect could not report one. The parakeet-cpp backend could therefore not refuse an encoder swap for them (ECAPA and CAM++ both give 192 values) and only had the file-name rule. They are now fingerprinted like voices enrolled from speaker_profiles.

  • Pin bump: voice-detect.cpp to cf9e1d5 (merged upstream; descends from the previous pin and is the head of master). It adds voicedetect_capi_encoder_arch, _name and _family; the C ABI version stays 1. The bump_deps workflow reads the same Makefile variable, nothing else to change.
  • Backend: the three symbols are bound with a Dlsym probe, so a libvoicedetect without them still loads and reports nothing. The returned strings are borrowed from the context: they are copied at once and never freed. NULL, empty or an all-empty family (:::) means unavailable. The model file is hashed once at load (streamed, sha256:<hex>). A model that is not a plain file leaves the weights empty (the backend only loads a GGUF path, so there is no bundle component case here), and the family alone fingerprints the voice.
  • Proto: VoiceEmbedResponse gets encoder_family = 3 and encoder_weights = 4 (strings, empty when unknown). Old backends never set them and old clients ignore them.
  • Registry: /v1/voice/register stores the family as encoder_family (existing field) and the weights in a new optional encoder_weights field. The existing model field is not reused for the hash here: /v1/voice/identify treats a sha256: model as a portable voice and asks the backend for trusted encoder metadata, which the voice-detect backend does not expose, so every such voice would be filtered out. model keeps the encoder name, so the name filter keeps working. The diarization request path sends the weights as encoder_weights in KnownVoice.
  • Selection: a voice that carries a family is sent to the backend whatever its file name, so the backend decides by family (match, or a clear error that names both families). Voices without a family keep the file-name filter.
  • Identify: when the stored voice and the probe both have a family, the family decides; otherwise the old name rule applies. Verify is untouched.
  • Docs: docs/content/features/voice-recognition.md describes the fingerprint, the mismatch behaviour, speaker_strict, re-registering old voices, and the conversion-path caveat.

Notes for Reviewers

Limits:

  • For CAM++, WeSpeaker and ERes2Net the GGUF general.name is a conversion path, so the family changes when the same encoder is converted again under another name. Voices registered with the first file are then refused by the second.
  • Two fine-tunes with the same architecture, name and size have the same family. Only the weights hash tells them apart, and a hash mismatch alone only warns (parakeet.cpp behaviour).
  • Voices registered before this change have no fingerprint and stay as they were (file-name rule, rejected by speaker_strict). They should be registered again. A voice from an older libvoicedetect has weights but no family: parakeet-cpp gives it the loaded encoder's family when the weights are the same file.

Tests:

  • New Ginkgo specs with stubbed C calls: symbols present, missing (nil), NULL pointer, empty string, :::, a family with an empty field (colons kept), NULL context, file hash for a file, a missing path and a directory, and the VoiceEmbed response fields with and without a fingerprint.
  • Registry and selection: metadata carries family and weights, old JSON entries load without them and omit them when written, an audio voice with a family is forwarded under another file name, an old voice without one keeps the name filter.
  • There is no core-level spec that runs register and diarization together (it needs both real backends); the smoke test below covers it.
  • go test ./backend/go/voice-detect/... ./backend/go/parakeet-cpp/... ./core/services/voicerecognition/... ./core/http/endpoints/localai/... ./core/backend/... pass; golangci-lint on the touched packages reports 0 issues. go vet on voice-detect shows the two unsafe.Pointer notes that were already there.

Smoke test on CPU, with real builds of the voice-detect backend at cf9e1d5, the parakeet-cpp backend at the master pin and local-ai from this branch. The models were the ECAPA, CAM++ and Nemotron-3-Diarization q8_0 files from the gallery entries (checksums match), configured by hand rather than installed from the gallery.

  • Two voices cut from a two-speaker fixture by the diarization turn times, registered with ECAPA through /v1/voice/register. The backend logged family voicedetect:ecapa_tdnn:speechbrain/spkrec-ecapa-voxceleb:192 and the file's sha256.
  • Diarization with speaker_model set to the same ECAPA file: all five segments named (scores 0.96 and 0.98).
  • Diarization with the CAM++ file (also 192 values): fails with speaker registry was made with encoder family voicedetect:ecapa_tdnn:...:192, but the encoder in use is voicedetect:campplus:...:192; the embeddings are not comparable, enroll again with this encoder. Without the fingerprint these voices would have been dropped silently by the name filter, so this also shows the voices carry the family.
  • /v1/voice/identify with ECAPA finds the registered voice (distance about 0); with CAM++ it returns no match. /v1/voice/verify works.
  • Symbol-missing fallback: the same backend with libvoicedetect built at the previous pin (no encoder_* symbols). It loads, reports an empty family and the file hash, registers, and diarization names both voices (weights equal, so the voice takes the loaded encoder's family).

Not verified

  • GPU builds and runs, non-Linux platforms, real image builds (only local backend packages on Linux).
  • A voice registered by a build before this change (no family and no weights) with speaker_strict:true: covered by the selection specs, not run against real backends. In the run with the older library the voices still had weights, so strict mode accepted them.
  • A weights-only mismatch (same family, another quantization) was not run end to end; that warning comes from parakeet.cpp and is covered by its own tests.
  • Distributed mode.

Signed commits

  • Yes, I signed my commits.
  • Documentation updated (docs/content/) for user-facing changes, or not applicable

🤖 Generated with Claude Code

// when the path is not a readable regular file; the backend then reports no
// weights identity and only the family fingerprints the voice.
func fileIdentity(path string) string {
f, err := os.Open(path)
@mudler
mudler force-pushed the feat/voice-detect-encoder-family branch from 989b44d to b2d91d9 Compare October 8, 2026 17:15
mudler added 3 commits October 8, 2026 17:15
…ces are fingerprinted

A voice registered from audio through the voice-detect backend had no
encoder fingerprint, so the parakeet-cpp backend could not tell whether
it was comparable with the loaded speaker model and could only fall back
to the file-name rule.

The voice-detect backend now binds the three new libvoicedetect accessors
with a symbol probe (an older library still loads and reports nothing),
copies the borrowed strings at once and never frees them. It also hashes
the model file once at load. VoiceEmbedResponse gains two optional
fields, encoder_family and encoder_weights ("sha256:<hex>", empty when
the model is not a plain file).

/v1/voice/register stores them as encoder_family and a new
encoder_weights field in the registry entry; model keeps the encoder
name, so the 1:N identify filter by name is unchanged for old entries.
A voice with a family is sent to the backend whatever its file name, and
the backend decides by family. /v1/voice/identify compares the family
when both the stored voice and the probe have one. Old entries load
without the fields and stay unfingerprinted.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
…ated

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
@mudler
mudler force-pushed the feat/voice-detect-encoder-family branch from b2d91d9 to 674ff38 Compare October 8, 2026 17:15
@mudler
mudler merged commit 6343a2d into master Oct 8, 2026
14 of 23 checks passed
@mudler
mudler deleted the feat/voice-detect-encoder-family branch October 8, 2026 17:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants