Embedded vs remote
Embedded is the default: inference on the database host. Remote mode runs inference on postvec-server, a separate process from PostgreSQL, multi-threaded on the same VM, serving a CPU or GPU fleet and managing models from the dashboard. Managed cloud databases use it too. When to use postvec-server.
postvec.mode is cluster-wide and SIGHUP: the worker reads it at start. postvec setup restarts the cluster on a mode change, because the launcher fixes its own mode when it starts. SQL, the job queue, retry policy and the gRPC wire contract stay the same.
| Embedded | Remote (grpc) with postvec-server | |
|---|---|---|
| Inference | One engine inside the PostgreSQL launcher (one thread) | postvec-server, multi-threaded, CPU or GPU |
| Discovery | Loopback HTTP on the launcher | GET /config on those servers |
| DB-host assets | Extension, CLI, ONNX Runtime, models | Extension + CLI |
| Raw text leaves the DB host | Stays on the host, except columns bound to an external provider | Goes to the configured servers. Provider-bound columns continue to the hosted provider |
| Model commands | postvec model pull / upgrade / rm / activate | On each server: CLI or dashboard |
| Model fault | Restarts the launcher | Stays in the postvec-server process |
| Dashboard | Port 22222 on the server | |
| Typical use | Single-node, private, edge, air-gapped | Isolation, GPU, a fleet, managed PostgreSQL |
The same package supports both modes. setup --embedded (the default) selects embedded; setup --grpc / --http selects remote.
Embedded
sudo postvec setup --database app --embedded
postvec model ls
sudo postvec doctor --database app --deepThe engine root must contain:
{root}/libs/**/libonnxruntime.so # unversioned filename required
{root}/models/{backend}/{name}/ # descriptor and weightsWorkers keep jobs pending until the engine listener is ready. model pull installs files deactivated; model activate is what loads them.
The embedded gRPC/HTTP listeners stay on 127.0.0.1.
Local inference, weights and text remain on the database host. A column bound to a hosted provider sends its source text to that provider from the launcher.
Remote gRPC
sudo postvec setup --database app \
--grpc 10.0.0.20:33333 \
--http https://10.0.0.20:22222On each inference host, the same command on every node:
postvec-server \
--peers node-1,node-2,node-3 \
--ssl-cert /etc/postvec-server/tls.crt \
--ssl-cert-key /etc/postvec-server/tls.keyEvery node must carry the same enabled models. postvec round-robins the endpoints it was given, so a converter present on some nodes and not others fails intermittently. postvec-server status --fleet is the check.
postvec-server is a CPU or GPU inference process on the deployment network. Every member of a fleet runs the same command and binds the same ports; models are files on disk that each process loads. The server also serves a dashboard on the discovery port. Install is one process; fleet is several.
The gRPC port is plaintext and unauthenticated by design, so the nodes belong on a trusted private network. External provider credentials, when required, stay on that side. Every node must carry the same connector files and the same models. A UniVec hosted converter sends stored vectors from the node to UniVec. --allow-unreachable is only for staging configuration before the nodes exist.
If the remote engine is unavailable, jobs remain pending without consuming retry attempts. Lexical search can still run while postvec.search_degrade_to_fts is on (the default).
Switching
Mode is cluster-wide. Switching restarts the cluster and requires every configured database name and the --switch-mode flag:
sudo postvec setup --database app \
--embedded --switch-modeEngine files already on the database host stay in place in remote mode.
Managed PostgreSQL
When the database cannot load postvec.so, the same SQL surface is installed as plain SQL and the worker runs inside postvec-server. See managed PostgreSQL.