Skip to content

Optional BM25 FTS ranking via pg_textsearch - #285

Draft
samuelvkwong wants to merge 3 commits into
hybrid-fused-result-cachefrom
pgtextsearch-bm25-ranking
Draft

samuelvkwong wants to merge 3 commits into
hybrid-fused-result-cachefrom
pgtextsearch-bm25-ranking

Conversation

@samuelvkwong

Copy link
Copy Markdown
Member

Draft, stacked on #284 (which stacks on #282). Follows the direction @medihack pointed at in #64 (comment) — pg_textsearch over ParadeDB — and is meant to make that evaluation concrete. Opt-in and inert by default.

Motivation (measured)

The FTS half of hybrid search ranks every match before its LIMIT applies. On a 1.7M-report corpus, "fraktur" matches 414,572 reports; one fusion costs 11–12 s (6.8 s of it ts_rank + sort, plan: parallel seq scan + top-N heapsort), per page. ts_rank also carries no corpus-level IDF — a match on a term appearing in 414k documents weighs the same as one appearing in 12. Both problems grow with the corpus; #284 amortizes the cost across pages, this PR is the structural fix: BM25 ranking out of an index (Block-Max WAND top-k), with IDF.

Why pg_textsearch and not PR #64's ParadeDB

Per the maintainer's comments there: paradedb#1793 (per-record language stemming) blocks ParadeDB — its per-index stemmer forces languages into schema (generated columns). pg_textsearch is PostgreSQL-licensed, builds on Postgres's own text search configs, and its documented multilingual pattern — one partial BM25 index per language — keeps language as data: this PR manages the indexes from the Language table at runtime (sync_bm25_indexes), not in migrations. Adding a language to a corpus means re-running a command, not writing a migration.

What's here

  • docker/postgres/Dockerfile — the pinned pgvector image + pg_textsearch v1.4.0 built from source (plain C extension), shared_preload_libraries set (it refuses to load otherwise — note: adopting this means one postgres restart).
  • HYBRID_FTS_RANKING = ts_rank | bm25 (default ts_rank — nothing changes until a deployment opts in; participates in Cache the fused RRF union per query fingerprint #284's cache key).
  • manage.py sync_bm25_indexes — creates the extension and one partial index per Language row (text_config via the existing code_to_language mapping).
  • Provider: in bm25 mode the boolean tsquery match still gates membership — AND/OR/NOT and phrase semantics are unchanged — while body <@> to_bm25query(...) provides the ordering. Per-language ranking (BM25 statistics live per index), merged by score with the caveat documented (mixed-corpus scores are approximate — as are cross-stemmer ts_rank values today).
  • Database checks pgsearch.E004/E005: bm25 selected but extension/indexes missing fails at migrate (init container), not on the first search.
  • Tests that skip on a stock pgvector image, so CI stays green without the extension.

Verified

  • pg_textsearch 1.4.0 and vector 0.8.6 coexist in one database on the built image.
  • German partial index builds; BM25 behaves (term-frequency saturation visible; non-matching docs score 0 — confirming membership must come from the filter, as implemented).
  • 4 new tests green against the BM25 image; full pgsearch + search suites green on both images; lint clean.

Open questions for review

  1. Query shape vs. index acceleration: with the membership filter and group join in the query, does the planner drive the BM25 index for top-k, or score per row? v1.4.0's release notes claim up-to-5× filtered top-k optimizations; a staging-scale benchmark (1.7M reports) is the follow-up before any default change.
  2. Maturity: the extension is young (open-sourced 2026). The default stays ts_rank until we're comfortable.
  3. sync_bm25_indexes builds take a table lock and scale with corpus size — acceptable as an operator-run step?
  4. Mixed-corpus score merging across per-language indexes (documented caveat) — acceptable, or rank-interleave instead?

🤖 Generated with Claude Code

https://claude.ai/code/session_01VDei6anDxfR5eoHhFfXBGs

@coderabbitai

coderabbitai Bot commented Aug 26, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

samuelvkwong and others added 3 commits August 26, 2026 14:16
The FTS half of hybrid search ranks every match before its LIMIT applies:
on a 1.7M-report corpus a common term ("fraktur", 414,572 matches) costs
11-12 s per fusion, and ts_rank carries no corpus-level IDF. This adds an
opt-in ranking mode backed by Timescale's pg_textsearch (PostgreSQL
license): membership still comes from the boolean tsquery match - AND/OR/
NOT and phrase semantics unchanged - while body <@> to_bm25query(...)
orders the matches with index-backed, IDF-aware BM25.

Follows the maintainer's direction on PR #64: ParadeDB is blocked by
paradedb#1793 because its per-index stemmer forces languages into schema;
pg_textsearch's documented multilingual pattern is one partial BM25 index
per language, which keeps language as data - sync_bm25_indexes manages the
indexes from the Language table at runtime, no migrations involved.

Ships a postgres Dockerfile (pinned pgvector image + pg_textsearch v1.4.0,
shared_preload_libraries set - the extension refuses to load otherwise;
verified coexisting with vector 0.8.6), database checks (pgsearch.E004/
E005) so a misconfigured deployment fails at migrate rather than on the
first search, and tests that skip on a stock pgvector image so CI stays
green without the extension. Default stays ts_rank; the mode participates
in the fused-cache key.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VDei6anDxfR5eoHhFfXBGs
One INFO line per fusion attributes the cost to its arms - the FTS ranking
query, the query embedding (or its cache hit), the HNSW beam pass, and the
in-process RRF - plus row counts and a degraded flag, so production logs
answer "where does a slow search spend its time" without extra tooling.
Cache hits log their own line with the serve time. The query is identified
by a short hash: search text can contain patient identifiers, so the log
correlates repeated searches without recording what was searched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VDei6anDxfR5eoHhFfXBGs

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant