Skip to content

BM25 ​

search() fuses two retrieval legs with Reciprocal Rank Fusion: a semantic leg, nearest neighbours in the vector column, and a lexical leg, BM25 over PostgreSQL text search (fts_score). Both legs run by default. When the query embedding fails and postvec.search_degrade_to_fts is on, search() runs the lexical leg alone. semantic_weight sets how much each leg's rank contributes to rrf_score.

Enable the GIN index ​

BM25 scoring is part of search(). A GIN index keeps the lexical leg fast on a large table. Pass create_fts_index => true at enable() or adopt():

sql
SELECT postvec.enable(
  'public.docs', 'body',
  model => 'sentence-transformers-all-minilm-l6-v2',
  create_fts_index => true,
  fts_config => 'pg_catalog.english'
);

Expected: a GIN index named postvec_fts_<registry_id> on to_tsvector(fts_config, source). Chunked entries index chunk_text on the destination.

fts_config selects the text-search configuration (stemming, stopwords). Default is pg_catalog.english. The value is stored on the entry and used for both indexing and query parsing.

sql
-- Hybrid (default): 0.5 vector, 0.5 BM25
SELECT d.id, d.body, s.rrf_score, s.semantic_rank, s.fts_rank, s.fts_score
  FROM postvec.search('public.docs', 'body', 'reset password') AS s
  JOIN public.docs AS d ON d.id = s.pk_value::bigint
 ORDER BY s.rrf_score DESC;

-- Keyword only
SELECT * FROM postvec.search(
  'public.docs', 'body', 'reset password',
  semantic_weight => 0.0
);

-- Vector only
SELECT * FROM postvec.search(
  'public.docs', 'body', 'reset password',
  semantic_weight => 1.0
);

fts_rank is the 1-based position in the lexical pool. fts_score is the BM25 score once corpus stats exist. Until the first successful refresh it is ts_rank_cd.

Matching uses websearch_to_tsquery: quoted phrases, OR and -exclusions have the usual PostgreSQL websearch meaning. BM25 then scores the positive query lexemes. A query with only negative terms ties at score 0.

With search_with_vector(), query_text feeds the lexical leg. Leave it at the default '' for a vector-only call.

Corpus statistics ​

BM25 needs document count N, average document length avgdl and per-term document frequency. The worker rebuilds those after writes, throttled to at least 30 seconds or ten times the last pass, whichever is longer. A refresh is one to_tsvector pass over the text. It runs inside the worker and delays embedding for that long.

After a bulk load, rebuild now (table owner):

sql
SELECT postvec.refresh_lexical_stats('public.docs', 'body');

A failed call raises and rolls back. A failed background refresh keeps the previous good statistics and records status().lexical_error. The worker retries after ten minutes.

sql
SELECT relation, lexical_docs, lexical_stats_age_seconds, lexical_error
  FROM postvec.status();

Automatic change detection uses track_counts = on. Vector updates also count, so a backfill can cause extra, still-throttled refreshes. TRUNCATE on the registered source invalidates statistics immediately and wakes the worker. After truncating a single partition, changing partition membership or changing text-search dictionaries, refresh manually when ranks need to follow at once.

Row-level security ​

On a source with RLS, search uses ts_rank_cd for roles that see only a subset of rows. Table owner (unless FORCE RLS), BYPASSRLS and superuser still get BM25. Ordinary readers can read status() lexical fields only when they would get BM25.

Change or remove ​

GoalAction
Rank by vectors only, this querysemantic_weight => 1.0
Rank by keywords only, this querysemantic_weight => 0.0
Drop the GIN, keep the entryDROP INDEX postvec_fts_<registry_id> (registry_id from status())
Change fts_config or rebuild the GINdisable() then enable() / adopt() again
Tear down the entrydisable(); a postvec-created GIN is dropped with it

create_fts_index and fts_config are fixed after the first enable() / adopt().

Flags ​

FlagWhereDefaultMeaning
create_fts_indexenable(), adopt()falseBuild the GIN at registration
fts_configenable(), adopt()pg_catalog.englishText-search configuration for index and query
semantic_weightsearch(), search_with_vector()0.51.0 vector, 0.0 BM25, in between hybrid
query_textsearch_with_vector()''Text for the lexical leg
search_degrade_to_ftsGUC, USERSETonIf query embedding fails, continue on the lexical leg

Scoring uses Lucene IDF with k1 = 1.2 and b = 0.75. Tokenization is PostgreSQL to_tsvector / websearch_to_tsquery under fts_config. Term frequencies use stored tsvector positions, at most 256 per lexeme, positions capped at 16383. Statistics cover non-NULL documents (chunks for recursive entries), including empty and stopword-only text. Metadata filters apply to candidate rows; corpus stats stay those of the full registered text.

The lexical leg is SQL in both modes. Query embedding runs inside the PostgreSQL process by default, or on postvec-server in remote mode.