Repository navigation
Optional BM25 FTS ranking via pg_textsearch - #285
Draft
samuelvkwong wants to merge 3 commits into
Draft
samuelvkwong wants to merge 3 commits into
samuelvkwong wants to merge 3 commits into
Conversation
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueThanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
The FTS half of hybrid search ranks every match before its LIMIT applies:
on a 1.7M-report corpus a common term ("fraktur", 414,572 matches) costs
11-12 s per fusion, and ts_rank carries no corpus-level IDF. This adds an
opt-in ranking mode backed by Timescale's pg_textsearch (PostgreSQL
license): membership still comes from the boolean tsquery match - AND/OR/
NOT and phrase semantics unchanged - while body <@> to_bm25query(...)
orders the matches with index-backed, IDF-aware BM25.
Follows the maintainer's direction on PR #64: ParadeDB is blocked by
paradedb#1793 because its per-index stemmer forces languages into schema;
pg_textsearch's documented multilingual pattern is one partial BM25 index
per language, which keeps language as data - sync_bm25_indexes manages the
indexes from the Language table at runtime, no migrations involved.
Ships a postgres Dockerfile (pinned pgvector image + pg_textsearch v1.4.0,
shared_preload_libraries set - the extension refuses to load otherwise;
verified coexisting with vector 0.8.6), database checks (pgsearch.E004/
E005) so a misconfigured deployment fails at migrate rather than on the
first search, and tests that skip on a stock pgvector image so CI stays
green without the extension. Default stays ts_rank; the mode participates
in the fused-cache key.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VDei6anDxfR5eoHhFfXBGs
One INFO line per fusion attributes the cost to its arms - the FTS ranking query, the query embedding (or its cache hit), the HNSW beam pass, and the in-process RRF - plus row counts and a degraded flag, so production logs answer "where does a slow search spend its time" without extra tooling. Cache hits log their own line with the serve time. The query is identified by a short hash: search text can contain patient identifiers, so the log correlates repeated searches without recording what was searched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VDei6anDxfR5eoHhFfXBGs
# Conflicts: # radis/pgsearch/providers.py
This was referenced Aug 27, 2026
samuelvkwong
force-pushed
the
hybrid-fused-result-cache
branch
from
September 22, 2026 07:52
0246776 to
0ee271d
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation (measured)
The FTS half of hybrid search ranks every match before its LIMIT applies. On a 1.7M-report corpus, "fraktur" matches 414,572 reports; one fusion costs 11–12 s (6.8 s of it
ts_rank+ sort, plan: parallel seq scan + top-N heapsort), per page.ts_rankalso carries no corpus-level IDF — a match on a term appearing in 414k documents weighs the same as one appearing in 12. Both problems grow with the corpus; #284 amortizes the cost across pages, this PR is the structural fix: BM25 ranking out of an index (Block-Max WAND top-k), with IDF.Why pg_textsearch and not PR #64's ParadeDB
Per the maintainer's comments there: paradedb#1793 (per-record language stemming) blocks ParadeDB — its per-index stemmer forces languages into schema (generated columns). pg_textsearch is PostgreSQL-licensed, builds on Postgres's own text search configs, and its documented multilingual pattern — one partial BM25 index per language — keeps language as data: this PR manages the indexes from the
Languagetable at runtime (sync_bm25_indexes), not in migrations. Adding a language to a corpus means re-running a command, not writing a migration.What's here
docker/postgres/Dockerfile— the pinned pgvector image + pg_textsearch v1.4.0 built from source (plain C extension),shared_preload_librariesset (it refuses to load otherwise — note: adopting this means one postgres restart).HYBRID_FTS_RANKING = ts_rank | bm25(default ts_rank — nothing changes until a deployment opts in; participates in Cache the fused RRF union per query fingerprint #284's cache key).manage.py sync_bm25_indexes— creates the extension and one partial index perLanguagerow (text_configvia the existingcode_to_languagemapping).body <@> to_bm25query(...)provides the ordering. Per-language ranking (BM25 statistics live per index), merged by score with the caveat documented (mixed-corpus scores are approximate — as are cross-stemmerts_rankvalues today).pgsearch.E004/E005: bm25 selected but extension/indexes missing fails atmigrate(init container), not on the first search.Verified
pgsearch+searchsuites green on both images; lint clean.Open questions for review
ts_rankuntil we're comfortable.sync_bm25_indexesbuilds take a table lock and scale with corpus size — acceptable as an operator-run step?🤖 Generated with Claude Code
https://claude.ai/code/session_01VDei6anDxfR5eoHhFfXBGs