Connector files
One TOML file per provider.
Copy-paste setup: OpenAI, Cohere, Amazon Bedrock, Gemini, Mistral, OpenRouter, UniVec. Shared rules: external providers.
File format
# /etc/postvec/providers.d/openai.toml
provider = "openai" # openai | openrouter | mistral | google | cohere | aws | univec
enabled = true # false parses and serves nothing
api_key_file = "/etc/postvec/keys/openai.key"
base_url = "https://api.openai.com" # optional origin
max_concurrent = 4 # outbound requests in flight
timeout_ms = 20000 # per HTTP attempt
[[models]]
name = "openai-text-embedding-3-small" # route name
provider_model_id = "text-embedding-3-small" # API id
dim = 1536 # responses are checked against this
max_batch = 512 # host sub-batches to this
max_tokens = 8191 # advertised as sequence_len
# space = "openai-text-embedding-3-small" # vector space; default is name
# priority = 1 # lower wins; omit for derived order
# added = "2026-09-08T00:00:00Z" # RFC 3339; unstamped routes precede stamped onesbase_url replaces the origin only. Each connector appends its own path. OpenAI-shaped APIs and UniVec embeds use /v1/embeddings. UniVec conversion uses /v1/convert, Cohere uses /v2/embed and Gemini uses its batch path. The URL must be absolute http or https with a host. It must not carry a query string, fragment, userinfo or trailing /.
A reverse proxy, a gateway, a self-hosted OpenAI-compatible endpoint or Azure's v1 API under /openai/v1 all work as a base_url. Azure OpenAI's classic API needs a deployment path and an api-version query, which this schema has no field for.
Plain http to anything but loopback is refused unless the file opts in:
base_url = "http://vllm.internal:8000"
allow_insecure_transport = true # required, and never the defaultA file that opts in still warns on every doctor run. Loopback needs no opt-in.
Referenced secret paths must be absolute. The CLI, the embedded launcher and postvec-server all resolve them and none share a working directory.
Key sources
Exactly one of three, mutually exclusive:
| Field | Use it when | Note |
|---|---|---|
api_key_file = "/path" | The usual choice | An absolute path to a regular file (not a symlink, not hard-linked) at 0600. A trailing newline is trimmed |
api_key_env = "OPENAI_API_KEY" | Containers, systemd units | Resolved in the inference process's environment. The postmaster's or the unit's, not your shell's |
api_key = "sk-..." | Last resort | The value sits in the 0600 file |
api_key_file is opened with O_NOFOLLOW and the mode is checked on the descriptor the host actually opens. Kubernetes projected secret volumes are symlinks into a ..data directory and are mounted world-readable, so they cannot be referenced this way. Use api_key_env there.
AWS Bedrock
AWS replaces api_key* with region plus either a Bedrock bearer token (bearer_token, bearer_token_file, bearer_token_env) or a static SigV4 pair (access_key_id and secret_access_key, each with its own _file / _env variant):
provider = "aws"
region = "us-east-1"
bearer_token_file = "/etc/postvec/keys/bedrock.key"
[[models]]
name = "aws-titan-embed-text-v2-0"
provider_model_id = "amazon.titan-embed-text-v2:0"
dim = 1024
max_batch = 1 # Titan invokes one text per requestThe signer takes static credentials only: no session tokens, no instance profile, no IMDS.
A few AWS specifics:
- The connector speaks the Amazon Titan embedding schema. Other vendors hosted on Bedrock use different body shapes and are not in the built-in catalogue.
regionis the whole endpoint. The connector buildsbedrock-runtime.<region>.amazonaws.comfrom it, sobase_urldoes not apply toawsandprovider addrefuses the flag. A region is restricted to[a-z0-9-].- Bedrock's
InvokeModelsends one text per request, somax_batchis- On a Titan-backed column, lower
postvec.batch_sizerather than raising the timeout.
- On a Titan-backed column, lower
provider addandprovider testprobe the bearer-token variant. For a SigV4 file, pass--no-verifyand confirm withprovider lsplus a first write.
Gemini and Cohere dimensions
Where a provider documents it, dim is also requested: Gemini receives outputDimensionality and Cohere v4 output_dimension.
- Cohere v4 produces 256, 512, 1024 or 1536 and nothing else. A descriptor asking for another width is refused at load.
- Cohere v3 widths are fixed per model (
embed-english-v3.0andembed-multilingual-v3.0are 1024, the-light-pair 384) and are checked exactly. - Gemini ships one model:
gemini-embedding-001, withdimanywhere in Google's documented128..=3072. Any other Gemini model id is refused at load. A reduced Gemini vector is not normalised by Google, so postvec renormalises it.
Providers that separate query text from stored text send search_query for search() and search_document for worker writes, embed() and migration re-embeds. Cohere, Gemini and UniVec use this distinction.
UniVec converter entries
UniVec is the only connector that accepts kind = "convert". An embed entry uses the common schema above. A converter adds the source and target names in the provider's vocabulary and in postvec's vocabulary:
provider = "univec"
api_key_file = "/etc/postvec/keys/univec.key"
[[models]]
name = "univec-convert-snowflake-to-bge-m3"
kind = "convert"
provider_source_id = "snowflake-arctic-embed-l-v2.0"
provider_model_id = "baai-bge-m3"
source_model = "snowflake-arctic-embed-l-v2.0"
target_model = "baai-bge-m3"
source_dim = 1024
dim = 1024
max_batch = 96provider_source_id and provider_model_id go on the UniVec request. source_model and target_model are matched against a postvec column and the requested migration target. source_dim validates every input vector; dim validates every output vector.
Converter entries refuse max_tokens. They also refuse a missing name or dimension, identical source and target names, or a connector type other than univec. One invalid entry prevents the whole file from loading.
Hosted converters are direct routes for migrate() and convert(). Embed-bridge resolution uses local models. To search a retired space without migrating, see search a retired space. UniVec hosted models has the hosted conversion sequence.
Loading rules
*.tomlin the directory, read in lexicographic order. Dot-files skipped.- A file that fails to parse, or whose key cannot be resolved, is skipped with a logged error. Every other provider and every local model keeps working.
- A file is checked as a whole. An unknown
providertype, a connector with no credential, anawsfile with noregion, an implausibledimor an unusablebase_urlskips the entire file, including its other models. - The schema is per connector. Unknown fields for that connector type are refused.
enabled = falseparses and serves nothing.- A public name may appear once. Two entries in one file: the file is refused. Two files claiming the same name: neither serves until one drops it.
- A name that a loaded local engine model already serves refuses the whole file at load. The local model keeps serving under that name.
- The directory is bounded as a whole: at most 32 connector files, 512 provider models and 256 total
max_concurrentacross every file. A single file holds at most 256[[models]]entries (one whole UniVec catalogue with headroom), is at most 256 KiB, and a referenced secret at most 16 KiB. Breaking a directory-wide ceiling fails the whole scan.
provider add and provider rm validate the directory as it would be after the write. A file the host would skip as a whole is refused before it is written, so existing models on that provider stay up.
The directory
postvec setup --embedded creates /etc/postvec/providers.d (0700, owned by the cluster owner), and so does the first provider add. The CLI creates it, because only the CLI knows which account owns the cluster. An absent directory is the default.
On a postvec-server node the directory is <server-root>/providers.d, and a key copied from postvec login lands in <server-root>/keys/<stem>.key. Run postvec provider add --path <server-root> on the node to write the file there.
The directory is owned by the account that reads it (or by root) and must not be group- or world-writable. The same account is the only one besides root that can rewrite the directory or any ancestor. Anyone who can write there can drop in a connector file and choose where this host sends source text or stored vectors. The serving host refuses those cases. provider add also refuses to write there, and doctor fails provider.directory. Read and execute bits only disclose which providers are configured; those stay a warning.
provider add and provider rm take an advisory lock on the directory for the whole read-modify-write, so two administrators running them at once cannot lose one another's change.
Change the location with postvec.providers_path (SIGHUP) or setup --embedded --providers-path DIR, and on a postvec-server node with --providers-path or POSTVEC_PROVIDERS_PATH.
postvec uninstall reports connector files and leaves them. Package removal does the same.
Model names
The built-in catalogue maps each --model id to a public route name and a vector space. same means the space equals the route name.
| Provider | Model id | Route name | Space | Dim |
|---|---|---|---|---|
| openai | text-embedding-3-small | openai-text-embedding-3-small | same | 1536 |
| openai | text-embedding-3-large | openai-text-embedding-3-large | same | 3072 |
| openai | text-embedding-ada-002 | openai-text-embedding-ada-002 | same | 1536 |
gemini-embedding-001 | gemini-embedding-001 | same | 3072 | |
| cohere | embed-v4.0 | cohere-embed-v4-0 | same | 1536 |
| cohere | embed-english-v3.0 | cohere-embed-english-v3-0 | same | 1024 |
| cohere | embed-multilingual-v3.0 | cohere-embed-multilingual-v3-0 | same | 1024 |
| aws | amazon.titan-embed-text-v2:0 | aws-titan-embed-text-v2-0 | same | 1024 |
| aws | amazon.titan-embed-text-v1 | aws-titan-embed-text-v1 | same | 1536 |
| mistral | mistral-embed | mistral-mistral-embed | same | 1024 |
| openrouter | openai/text-embedding-3-small | openrouter-openai-text-embedding-3-small | openai-text-embedding-3-small | 1536 |
| openrouter | openai/text-embedding-3-large | openrouter-openai-text-embedding-3-large | openai-text-embedding-3-large | 3072 |
| openrouter | openai/text-embedding-ada-002 | openrouter-openai-text-embedding-ada-002 | openai-text-embedding-ada-002 | 1536 |
| openrouter | google/gemini-embedding-001 | openrouter-google-gemini-embedding-001 | gemini-embedding-001 | 3072 |
| openrouter | cohere/embed-v4.0 | openrouter-cohere-embed-v4-0 | cohere-embed-v4-0 | 1536 |
| openrouter | cohere/embed-english-v3.0 | openrouter-cohere-embed-english-v3-0 | cohere-embed-english-v3-0 | 1024 |
| openrouter | cohere/embed-multilingual-v3.0 | openrouter-cohere-embed-multilingual-v3-0 | cohere-embed-multilingual-v3-0 | 1024 |
| openrouter | mistralai/mistral-embed | openrouter-mistralai-mistral-embed | mistral-mistral-embed | 1024 |
These 18 entries are the whole catalogue. An OpenRouter entry mirrors a vendor route, so it shares that route's space: a column bound to openai-text-embedding-3-small is served by openrouter-openai-text-embedding-3-small once only the OpenRouter key exists.
Any other id works too. The probe measures the dimension, or --dim states it. A space has one width: two embed entries in one file that disagree refuse the file; an entry that disagrees with another file or with a loaded local model is skipped at load time (the rest of its file serves), and the cache refresh skips a contradicting row with a WARNING.
The built-in catalogue covers the six connectors above, not UniVec. provider add univec --model baai-bge-m3 takes the dimension from UniVec's catalogue, or measures it for an id the catalogue lacks, and writes route univec-baai-bge-m3.
For an unlisted id, the name is derived mechanically: lowercase, and every character outside [a-z0-9._-] becomes -. provider_model_id keeps the id exactly as the API expects it.
gemini is accepted as an alias for google, and amazon for aws. The file always records the canonical name. An alias never changes what you type in SQL.
Commands
| Command | Result | Network |
|---|---|---|
provider add TYPE --model ID... | Write or extend the file, verify, reload the host, refresh postvec.models | One embed per verified model, unless --no-verify |
provider add univec --convert-to MODEL | Write or extend the file with every catalogue converter into MODEL, verify, reload and refresh | One vector conversion per added entry, unless --no-verify |
provider ls | Providers, key sources, models, dims and whether the host serves them now | Loopback |
provider test NAME [--model ID] | Verify embed or converter entries on demand | One embed or vector conversion per selected entry |
provider rm NAME [--model ID] | Drop one model entry or the whole file, then reload | Loopback |
Useful options on add:
| Option | Effect |
|---|---|
--name STEM | Write STEM.toml instead of the type's name. Two OpenAI-compatible endpoints, two files |
--api-key-file P / --api-key-env VAR / --key-stdin | Choose the key source. Without any of them and with a TTY, a hidden prompt asks and records the key inline as api_key |
--api-key-from-login | univec only. Copy the key postvec login stored into <providers root>/keys/<stem>.key and use it for inference. An interactive run offers this; a script must ask. Refused when the file holds a different key |
--replace-copied-key | With --api-key-from-login: rotate a key file this command copied for this connector earlier. The new key is verified before the old one is replaced |
--base-url URL | Gateways, Azure-shaped fronts or a mock server. Not accepted for aws |
--region | Required for aws, and only valid there |
--dim N | Vector dimension for a single model, and the requested width where the provider takes one (Gemini, Cohere v4). Required with --no-verify when the catalogue and the probe cannot supply it |
--space SPACE | Vector space this route serves. Default: the catalogue space, else the public name |
--prefer | Give the new route priority 1 in its space, ahead of the local model |
--convert SRC:DST | univec catalogue converter, by provider id pair. Repeatable |
--convert-to MODEL | univec: every catalogue converter into MODEL. Repeatable |
--convert-from MODEL | univec: every catalogue converter out of MODEL. Repeatable |
--all-converters | univec: every catalogue converter |
--no-catalog | univec: skip the catalogue fetch. Every dimension then comes from --dim or the probe |
--path DIR | Write into DIR/providers.d instead of the selected cluster. Default: POSTVEC_PROVIDERS_PATH, then the cluster's providers path |
--acknowledge-in-use | Accept that columns bound to these names start sending their text to the provider at the next worker cycle. Required with --yes when a column is affected |
--yes | Skip the confirmation prompt. Required for a mutation without a TTY |
--no-verify | Skip the live probe |
--dry-run | Print the plan and change nothing |
--convert-source ID --convert-target ID | Hidden. Manual converter: UniVec provider ids |
--source-model NAME --target-model NAME --source-dim N | Hidden. Manual converter: postvec route names and the input dimension |
--converter-name NAME | Override the derived univec-convert-<source>-to-<target> name |
The plan is confirmed before the probe. Declining it, or failing the in-use acknowledgement, makes no API call.
add and test use an embedding probe for embed entries and a single-vector conversion probe for converter entries. A connector change, such as a new key source or base_url, rechecks every entry in that file.
Target resolution:
| Target | Behaviour |
|---|---|
--path DIR | Filesystem management of DIR/providers.d, or of DIR itself when it already is one. DIR must already exist. New files inherit its owner. An embed entry requires --acknowledge-in-use; a new converter skips that flag |
| Embedded cluster | Manage postvec.providers_path, scan the databases for affected columns, reload the running host |
Remote fleet (grpc) | Refused by name on the database host, because connector files live on the inference nodes. The message points at --path; run provider add --path <server-root> on a postvec-server node |
Reading provider ls:
providers.d: /etc/postvec/providers.d
openai (openai, key: file:/etc/postvec/keys/openai.key)
openai-text-embedding-3-small dim 1536 served
cohere (cohere, key: env:COHERE_API_KEY)
cohere-embed-v4-0 dim 1536 NOT served (reload or restart the host)
azure (openai, key: env:AZURE_KEY) [enabled = false]
openai-text-embedding-3-large dim 3072 disabled in the file
mistral (mistral, key: none)
! the host REFUSES this file: provider "mistral" needs an API key: set
exactly one of api_key_file (recommended), api_key_env or api_key
mistral-mistral-embed dim 1024 not loadable (see above)NOT served after a hand-edit: reload or restart. enabled = false: reported as disabled. REFUSES: the file will not load. ls and doctor apply the loader's rules.
doctor checks
postvec doctor has a provider.* family, all read-only:
| Check | Reports |
|---|---|
provider.directory | Fails when the directory is group- or world-writable, or when ownership or ancestors fail the write-safety rules. Warns on mere read/execute bits. "Does not exist" is a pass |
provider.file | A file the serving host would refuse: mode, unknown fields, two sources for one secret, an unknown type, no credential, a bad dim / region / base_url, no [[models]]. Warns on a plaintext base_url to a non-loopback host |
provider.key-source | A referenced key file that is missing, a symlink or too permissive, or a named variable that is absent. Every source, including both halves of an AWS SigV4 pair |
provider.served | A configured model missing from the running host. Skipped for enabled = false files |
A complaint about an environment variable can be a false alarm. doctor observes its own environment, and the postmaster's is what matters. The check says so.
provider add / rm write the files first and then ask the host to reload. When no host answers on a cluster target, the result is partial (exit 3): the files are correct and a restart applies them. On a --path target the same situation is a note, not a partial result.
Limits
| Topic | What the product does |
|---|---|
| Credentials | Keys live in providers.d on the inference host. PostgreSQL holds a path |
| Spend control | max_concurrent bounds calls in flight. Set quotas on the provider |
| AWS auth | Static SigV4 pair or a Bedrock bearer token |
| Azure OpenAI | base_url fronts the /openai/v1 API |
| Gemini | gemini-embedding-001 only (the documented contract) |
| Hosted converters | Direct routes for migrate() and convert(). Embed-bridge uses local models |
| Provider TLS | rustls with the bundled Mozilla root set |
| Inference transport | Loopback gRPC (embedded) and postvec-server gRPC (remote) are plaintext. Restrict the gRPC port. Use provider-side quotas as the spend control |