Skip to content

Limits and compatibility ​

Supported ​

  • PostgreSQL 16, 17, 18
  • pgvector >= 0.8
  • Debian 12, Ubuntu 22.04 / 24.04, EL9 (Alma, Rocky; RHEL/CentOS Stream with --force-untested on the bootstrap)
  • amd64 and arm64
  • Ordinary and partitioned tables with a primary key

Environment ​

On a self-hosted cluster, postvec loads as shared_preload_libraries = 'postvec' and the host restarts PostgreSQL. Tables must be ordinary or partitioned and have a primary key. The worker runs in its own session, so TEMPORARY tables are invisible to it.

On RDS, Aurora, Cloud SQL, Azure Flexible Server, Supabase and Neon, managed PostgreSQL installs a plain SQL schema and runs the worker in postvec-server. The same binary is the remote inference process for self-hosted clusters (install).

Scope ​

postvec covers in-database embedding, hybrid search, chunking, adopt and in-place migration. Generation (rag(), chat completion) and document parsing (PDF, HTML, Office) stay in the application. Storage and ANN indexes are pgvector's. Provider API keys live in the inference layer; see external providers (OpenAI, Cohere, Amazon Bedrock, Gemini, Mistral, OpenRouter, UniVec).

The Bedrock signer takes static credentials. max_concurrent bounds provider calls in flight. Connectors use rustls with the bundled Mozilla root set. Hosted converters serve migrate() and convert(). Embed-bridge uses local models.

SQL constraints ​

SituationWhat happens
Chunked adopt()Use enable(..., chunking => 'recursive') for new chunked entries
Chunked index_mode => 'immediate'Index the destination after backfill
Composite PK + chunkingColumn mode accepts composite PKs; chunking needs a single-column PK
Write made directly to a partitionThe default trigger_mode => 'statement' covers writes through the parent only. Use trigger_mode => 'row', which clones the trigger to each partition
Live splitter reconfigurationdisable, drop, enable again
halfvec / undimensioned vector on adoptRewrite to vector(N) first; the error includes the ALTER TABLE
NOT NULL vector the worker would writeObserved adopt (sync => false, backfill => 'none') or drop the constraint
DROP EXTENSION ... CASCADEUse uninstall for bounded teardown

Resource notes ​

  • One worker per configured database. An auto index build occupies that worker.
  • Chunked writers pay invalidation inside their own transaction.
  • set_format() blocks writers while it enqueues a full refresh.
  • query_timeout_ms (default 2 s) bounds the synchronous query embed.
  • Chunk splitter: at most 10,000 non-blank chunks, 32 MiB UTF-8 and 4x amplification per document.
  • postvec.max_document_bytes (default 1 MiB) dead-letters a row whose rendered text is larger, with the measured size. Text is never truncated. The same ceiling bounds embed() and search() query input.
  • postvec.max_batch_total_bytes (default 16 MiB) ends a batch at that budget. Rows past it stay pending for the next cycle without consuming a retry attempt.
  • With postvec-server, the engine is a separate multi-threaded process, so an inference fault stays inside it.

License split ​

LayerLicense
Extension, CLI, packages, PostgreSQL imagesPostgreSQL License
Embedded inference on the database hostThe same stack, running in-process
postvec-server (remote-mode node and managed PostgreSQL)Business Source License 1.1
Converter catalogueUniVec, under a separate license
Bundled MiniLMUpstream license, shipped in the model package

A verified UniVec account sees the private catalogue superset. The public channel is a subset and needs no key. Terms and Pro: License.