BM25
search() fuses two retrieval legs with Reciprocal Rank Fusion: a semantic leg, nearest neighbours in the vector column, and a lexical leg, BM25 over PostgreSQL text search (fts_score). Both legs run by default. When the query embedding fails and postvec.search_degrade_to_fts is on, search() runs the lexical leg alone. semantic_weight sets how much each leg's rank contributes to rrf_score.
Enable the GIN index
BM25 scoring is part of search(). A GIN index keeps the lexical leg fast on a large table. Pass create_fts_index => true at enable() or adopt():
SELECT postvec.enable(
'public.docs', 'body',
model => 'sentence-transformers-all-minilm-l6-v2',
create_fts_index => true,
fts_config => 'pg_catalog.english'
);Expected: a GIN index named postvec_fts_<registry_id> on to_tsvector(fts_config, source). Chunked entries index chunk_text on the destination.
fts_config selects the text-search configuration (stemming, stopwords). Default is pg_catalog.english. The value is stored on the entry and used for both indexing and query parsing.
Search
-- Hybrid (default): 0.5 vector, 0.5 BM25
SELECT d.id, d.body, s.rrf_score, s.semantic_rank, s.fts_rank, s.fts_score
FROM postvec.search('public.docs', 'body', 'reset password') AS s
JOIN public.docs AS d ON d.id = s.pk_value::bigint
ORDER BY s.rrf_score DESC;
-- Keyword only
SELECT * FROM postvec.search(
'public.docs', 'body', 'reset password',
semantic_weight => 0.0
);
-- Vector only
SELECT * FROM postvec.search(
'public.docs', 'body', 'reset password',
semantic_weight => 1.0
);fts_rank is the 1-based position in the lexical pool. fts_score is the BM25 score once corpus stats exist. Until the first successful refresh it is ts_rank_cd.
Matching uses websearch_to_tsquery: quoted phrases, OR and -exclusions have the usual PostgreSQL websearch meaning. BM25 then scores the positive query lexemes. A query with only negative terms ties at score 0.
With search_with_vector(), query_text feeds the lexical leg. Leave it at the default '' for a vector-only call.
Corpus statistics
BM25 needs document count N, average document length avgdl and per-term document frequency. The worker rebuilds those after writes, throttled to at least 30 seconds or ten times the last pass, whichever is longer. A refresh is one to_tsvector pass over the text. It runs inside the worker and delays embedding for that long.
After a bulk load, rebuild now (table owner):
SELECT postvec.refresh_lexical_stats('public.docs', 'body');A failed call raises and rolls back. A failed background refresh keeps the previous good statistics and records status().lexical_error. The worker retries after ten minutes.
SELECT relation, lexical_docs, lexical_stats_age_seconds, lexical_error
FROM postvec.status();Automatic change detection uses track_counts = on. Vector updates also count, so a backfill can cause extra, still-throttled refreshes. TRUNCATE on the registered source invalidates statistics immediately and wakes the worker. After truncating a single partition, changing partition membership or changing text-search dictionaries, refresh manually when ranks need to follow at once.
Row-level security
On a source with RLS, search uses ts_rank_cd for roles that see only a subset of rows. Table owner (unless FORCE RLS), BYPASSRLS and superuser still get BM25. Ordinary readers can read status() lexical fields only when they would get BM25.
Change or remove
| Goal | Action |
|---|---|
| Rank by vectors only, this query | semantic_weight => 1.0 |
| Rank by keywords only, this query | semantic_weight => 0.0 |
| Drop the GIN, keep the entry | DROP INDEX postvec_fts_<registry_id> (registry_id from status()) |
Change fts_config or rebuild the GIN | disable() then enable() / adopt() again |
| Tear down the entry | disable(); a postvec-created GIN is dropped with it |
create_fts_index and fts_config are fixed after the first enable() / adopt().
Flags
| Flag | Where | Default | Meaning |
|---|---|---|---|
create_fts_index | enable(), adopt() | false | Build the GIN at registration |
fts_config | enable(), adopt() | pg_catalog.english | Text-search configuration for index and query |
semantic_weight | search(), search_with_vector() | 0.5 | 1.0 vector, 0.0 BM25, in between hybrid |
query_text | search_with_vector() | '' | Text for the lexical leg |
search_degrade_to_fts | GUC, USERSET | on | If query embedding fails, continue on the lexical leg |
Scoring uses Lucene IDF with k1 = 1.2 and b = 0.75. Tokenization is PostgreSQL to_tsvector / websearch_to_tsquery under fts_config. Term frequencies use stored tsvector positions, at most 256 per lexeme, positions capped at 16383. Statistics cover non-NULL documents (chunks for recursive entries), including empty and stopword-only text. Metadata filters apply to candidate rows; corpus stats stay those of the full registered text.
The lexical leg is SQL in both modes. Query embedding runs inside the PostgreSQL process by default, or on postvec-server in remote mode.