Skip to content

postvec-server reference ​

postvec-server --help is the authority. This page is for reading what a setting does and where it can be set.

Configuration precedence ​

Defaults < config file < environment < flags. A present flag always wins. An omitted flag leaves the file or environment value in place. No file is required. POSTVEC_PATH and POSTVEC_PROVIDERS_PATH are shared with the postvec CLI. Use absolute paths; empty values are ignored. The providers variable names the exact connector directory.

The file is searched at $root/postvec-server.json, /etc/postvec-server/config.json, /etc/postvec-server.json, then ./postvec-server.json, unless --config names one. In that case a missing file is an error.

Unknown keys are a hard error, and the message names the key. Keys beginning with // are comments.

FlagEnvironmentFile keyDefault
--root PATHPOSTVEC_PATH-/opt/postvec
--config PATHPOSTVEC_SERVER_CONFIG-The search order above
--bind ADDRPOSTVEC_SERVER_BINDbind_address0.0.0.0
--http PORTPOSTVEC_SERVER_HTTPhttp_port22222
--grpc PORTPOSTVEC_SERVER_GRPCgrpc_port33333
--gossip PORTPOSTVEC_SERVER_GOSSIPgossip_port11111
--admin PORTPOSTVEC_SERVER_ADMINadmin_port22223
--peers LISTPOSTVEC_SERVER_PEERS, then POSTVEC_SERVER_CLUSTERpeers (or cluster)Empty, a single node
--group NAMEPOSTVEC_SERVER_GROUPgrouppostvec
--advertise IPPOSTVEC_SERVER_ADVERTISEadvertise--bind if specific, else autodetected
--frontend URLPOSTVEC_SERVER_FRONTENDfrontend{scheme}://{advertise}:{http}
--ssl-cert PATHPOSTVEC_SERVER_SSL_CERTssl.cert$root/certs/server.crt
--ssl-cert-key PATHPOSTVEC_SERVER_SSL_KEYssl.key$root/certs/server.key
--insecurePOSTVEC_SERVER_INSECUREinsecureOff. TLS required
--models LISTPOSTVEC_SERVER_MODELSmodelsEmpty, every enabled descriptor
--providers-path PATHPOSTVEC_PROVIDERS_PATHproviders_path$root/providers.d. A path, never a credential
--predict-timeout-ms NPOSTVEC_SERVER_PREDICT_TIMEOUT_MSpredict_timeout_ms30000
--max-inflight NPOSTVEC_SERVER_MAX_INFLIGHTmax_inflightCPU count, clamped to 4-16
--max-resident-models NPOSTVEC_SERVER_MAX_RESIDENT_MODELSmax_resident_models16
--drain-delay-ms NPOSTVEC_SERVER_DRAIN_DELAY_MSdrain_delay_ms5000
--no-warmupPOSTVEC_SERVER_WARMUP=0warmupWarmup on
--no-metricsPOSTVEC_SERVER_METRICS=0metricsMetrics on
--log-level FILTERRUST_LOGlog_levelinfo
--managePOSTVEC_SERVER_MANAGE=1manage: trueOff. Also serve registry pull/activate/deactivate/remove on the public port, so the dashboard can manage models; see HTTP API
--web-ui DIRPOSTVEC_SERVER_WEB_UIweb_ui<root>/server/ui, where the package installs it

--cluster is an accepted alias for --peers, and --ssl-key for --ssl-cert-key. --remote is the deprecated spelling of --gossip; it is accepted and errors when the two disagree.

--max-inflight and --max-resident-models ​

--max-inflight bounds concurrently executing predictions. It is a CPU bound, because each ONNX session runs its own intra-op thread pool and N concurrent predictions can oversubscribe an N-core box several times over. It is also a memory bound, because each in-flight request may hold a response tree up to the transport envelope. Raise it on a large node when latency is fine and throughput is not. Lower it when the box is shared.

--max-resident-models bounds how many models the engine holds, counting whole dependency closures: a bridge model pulls its converters in with it. Resident models are the dominant memory cost on a node. Exceeding the ceiling is refused at boot, before any model loads.

frontend ​

--frontend sets the HTTP URL printed in /config, postvec-server status and membership. Use it when NAT, a load balancer or a DNS name should appear instead of https://{advertise}:{http_port}.

gRPC uses the advertised IP, membership uses the gossip socket, and postvec's discovery ignores frontend. Set --advertise (and --frontend when the printed URL should differ) to the address peers and clients actually use.

Health routes ​

RouteMeaning
GET /healthThe process is up. Stays 200 through a drain
GET /readyA model can answer right now. 503 before the first model loads, and for the whole drain
GET /configModel discovery, which is what postvec reads. Also carries server, cluster and system blocks
GET /metricsPrometheus text
GET /api/{model}Model layer overview, native envelope
POST /api/{model}Native inference. JSON body keyed like executor.inputs (texts / embeddings). Envelope {success, data} or {success, error:{message}}; HTTP stays 200
POST /api/convertThe SQL convert() over HTTP. Body {"source_model", "target_model", "embeddings"}; the node resolves the converter for the pair, and 404 names the pair when none serves it. See HTTP API
GET /api/registry/{models,available,pulls}, POST /api/registry/{pull,activate,deactivate,remove}Model management, the postvec model commands over HTTP. The POSTs are on the admin port, and on the public port with --manage. See HTTP API and dashboard
GET /admin/managed, GET /admin/managed/{name}/jobs, POST /admin/managed/{name}/refresh-models, POST /admin/managed/{name}/retry-deadManaged PostgreSQL state, recent jobs, model refresh and dead-letter retry. Loopback admin port only. See managed PostgreSQL
POST /api/openai/embeddingsOpenAI /v1/embeddings adaptor in front of the native path. {object, data, model, usage} on success; {error:{message, type, code}} and a real status on failure. See HTTP API

Gate load balancers and compose healthchecks on /ready, and supervisors on /health. Confusing the two produces either a node that receives traffic before it can serve it, or a supervisor that restarts a node mid-drain.

Boot warmup runs one request per embedding model, so /ready means the node answers at normal latency. --no-warmup skips it, and postvec_server_warmup_failures_total counts what went wrong.

Metrics ​

SeriesWhy
postvec_server_ready0 for longer than a restart takes means the node is not coming back
postvec_server_request_errors_total{code="MODEL_NOT_LOADED"}Almost always inventory drift across the fleet
postvec_server_request_errors_total{code="TIMEOUT"}Requests exceeding their budget. Raise --max-inflight, add nodes, or accept the latency
postvec_server_request_duration_secondsThe p99 your database's search() inherits
postvec_server_models_enabled_on_disk > ..._models_loadedSomething was pulled or activated and never loaded
postvec_server_cluster_membersBelow the fleet size means a partition or a wrong --advertise
postvec_managed_queue_depth{db}, postvec_managed_heartbeat_age_seconds{db}, postvec_managed_leader{db}Managed PostgreSQL worker state, one series per managed[] entry
postvec_managed_jobs_total{db,outcome}Managed worker outcomes (embedded, nulled, retried, dead). Database-wide and it survives a leader change, so scrape one node
postvec_proxy_connections{db}, postvec_proxy_rewrites_total{db,kind}Managed search proxy load; kind is search or embed

Error codes are the same values that cross the wire and drive postvec's retry and dead-letter policy.

Security ​

The gRPC and discovery ports are unauthenticated. gRPC is also plaintext. postvec's client speaks plaintext and the extension's own setting help says so, so adding TLS on the server side alone would break every existing postvec setup --grpc.

Restrict both ports to a private network with a firewall, a security group or --bind <private-ip>. The node logs a warning at boot whenever it binds every interface.

Load, unload, provider reload and managed recovery stay on the loopback admin port. /admin/load, /admin/unload, /admin/providers/reload and /admin/managed* bind 127.0.0.1 only. A routable admin bind is a boot failure. A per-request loopback peer check sits behind that. Registry pull/activate/deactivate/remove use the same admin socket; --manage also serves those POSTs on the public discovery port so the dashboard can drive them. postvec-server managed status prints the managed state as JSON, including worker_alive, the leader, the queue depth and the dead-letter count. The trust boundary is local OS users, plus whoever can reach port 22222 when --manage is on.

No telemetry and no licence check. Models arrive on disk by whatever mechanism you choose. An external provider connector is the exception: the node then calls that provider's API for the models it declares. A UniVec converter sends stored vectors; an embed entry sends text.

Troubleshooting ​

SymptomMeaning or next action
A node sees only itselfThe advertised address. Check the boot log's autodetection warning, pin --advertise, then open the gossip port
MODEL_NOT_LOADED, intermittentlyInventory drift. postvec-server status --fleet names the model and the nodes missing it
MODEL_NOT_LOADED, always, for a model you just pulledOn disk but not resident. postvec-server load NAME
Refuses to start with a TLS errorNo readable certificate pair and no --insecure
"cannot initialise ONNX Runtime"No libonnxruntime.so under $root/libs. Install postvec-onnxruntime or point --root at a tree that has one
A configuration change had no effectThe boot log's configuration file: line names the file that was read. Flags beat both the environment and the file
Refuses to start over a configuration keyUnknown keys are fatal by design. The error names the key; prefix it with // if you meant a comment
Requests queue and time out under loadpostvec_server_requests_in_flight sitting at --max-inflight with rising TIMEOUT. Raise it if the box has headroom, or add nodes. Lowering --predict-timeout-ms makes the failures faster, not fewer