Repository navigation
feat(voice-detect): report the encoder family so audio-registered voices are fingerprinted - #12499
Merged
Merged
Conversation
| // when the path is not a readable regular file; the backend then reports no | ||
| // weights identity and only the family fingerprints the voice. | ||
| func fileIdentity(path string) string { | ||
| f, err := os.Open(path) |
mudler
force-pushed
the
feat/voice-detect-encoder-family
branch
from
October 8, 2026 17:15
989b44d to
b2d91d9
Compare
…ces are fingerprinted
A voice registered from audio through the voice-detect backend had no
encoder fingerprint, so the parakeet-cpp backend could not tell whether
it was comparable with the loaded speaker model and could only fall back
to the file-name rule.
The voice-detect backend now binds the three new libvoicedetect accessors
with a symbol probe (an older library still loads and reports nothing),
copies the borrowed strings at once and never frees them. It also hashes
the model file once at load. VoiceEmbedResponse gains two optional
fields, encoder_family and encoder_weights ("sha256:<hex>", empty when
the model is not a plain file).
/v1/voice/register stores them as encoder_family and a new
encoder_weights field in the registry entry; model keeps the encoder
name, so the 1:N identify filter by name is unchanged for old entries.
A voice with a family is sent to the backend whatever its file name, and
the backend decides by family. /v1/voice/identify compares the family
when both the stored voice and the probe have one. Old entries load
without the fields and stay unfingerprinted.
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
…ated Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
mudler
force-pushed
the
feat/voice-detect-encoder-family
branch
from
October 8, 2026 17:15
b2d91d9 to
674ff38
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Voices registered from audio through the voice-detect backend (
POST /v1/voice/registerwithaudio) had no encoder fingerprint, because libvoicedetect could not report one. The parakeet-cpp backend could therefore not refuse an encoder swap for them (ECAPA and CAM++ both give 192 values) and only had the file-name rule. They are now fingerprinted like voices enrolled fromspeaker_profiles.master). It addsvoicedetect_capi_encoder_arch,_nameand_family; the C ABI version stays 1. Thebump_depsworkflow reads the same Makefile variable, nothing else to change.Dlsymprobe, so a libvoicedetect without them still loads and reports nothing. The returned strings are borrowed from the context: they are copied at once and never freed. NULL, empty or an all-empty family (:::) means unavailable. The model file is hashed once at load (streamed,sha256:<hex>). A model that is not a plain file leaves the weights empty (the backend only loads a GGUF path, so there is no bundle component case here), and the family alone fingerprints the voice.VoiceEmbedResponsegetsencoder_family = 3andencoder_weights = 4(strings, empty when unknown). Old backends never set them and old clients ignore them./v1/voice/registerstores the family asencoder_family(existing field) and the weights in a new optionalencoder_weightsfield. The existingmodelfield is not reused for the hash here:/v1/voice/identifytreats asha256:model as a portable voice and asks the backend for trusted encoder metadata, which the voice-detect backend does not expose, so every such voice would be filtered out.modelkeeps the encoder name, so the name filter keeps working. The diarization request path sends the weights asencoder_weightsinKnownVoice.docs/content/features/voice-recognition.mddescribes the fingerprint, the mismatch behaviour,speaker_strict, re-registering old voices, and the conversion-path caveat.Notes for Reviewers
Limits:
general.nameis a conversion path, so the family changes when the same encoder is converted again under another name. Voices registered with the first file are then refused by the second.speaker_strict). They should be registered again. A voice from an older libvoicedetect has weights but no family: parakeet-cpp gives it the loaded encoder's family when the weights are the same file.Tests:
:::, a family with an empty field (colons kept), NULL context, file hash for a file, a missing path and a directory, and theVoiceEmbedresponse fields with and without a fingerprint.go test ./backend/go/voice-detect/... ./backend/go/parakeet-cpp/... ./core/services/voicerecognition/... ./core/http/endpoints/localai/... ./core/backend/...pass;golangci-linton the touched packages reports 0 issues.go veton voice-detect shows the twounsafe.Pointernotes that were already there.Smoke test on CPU, with real builds of the voice-detect backend at cf9e1d5, the parakeet-cpp backend at the master pin and
local-aifrom this branch. The models were the ECAPA, CAM++ and Nemotron-3-Diarization q8_0 files from the gallery entries (checksums match), configured by hand rather than installed from the gallery./v1/voice/register. The backend logged familyvoicedetect:ecapa_tdnn:speechbrain/spkrec-ecapa-voxceleb:192and the file's sha256.speaker_modelset to the same ECAPA file: all five segments named (scores 0.96 and 0.98).speaker registry was made with encoder family voicedetect:ecapa_tdnn:...:192, but the encoder in use is voicedetect:campplus:...:192; the embeddings are not comparable, enroll again with this encoder. Without the fingerprint these voices would have been dropped silently by the name filter, so this also shows the voices carry the family./v1/voice/identifywith ECAPA finds the registered voice (distance about 0); with CAM++ it returns no match./v1/voice/verifyworks.encoder_*symbols). It loads, reports an empty family and the file hash, registers, and diarization names both voices (weights equal, so the voice takes the loaded encoder's family).Not verified
speaker_strict:true: covered by the selection specs, not run against real backends. In the run with the older library the voices still had weights, so strict mode accepted them.Signed commits
🤖 Generated with Claude Code