Skip to content

Connector files ​

One TOML file per provider.

Copy-paste setup: OpenAI, Cohere, Amazon Bedrock, Gemini, Mistral, OpenRouter, UniVec. Shared rules: external providers.

File format ​

toml
# /etc/postvec/providers.d/openai.toml
provider = "openai"      # openai | openrouter | mistral | google | cohere | aws | univec
enabled  = true          # false parses and serves nothing

api_key_file = "/etc/postvec/keys/openai.key"

base_url       = "https://api.openai.com"  # optional origin
max_concurrent = 4       # outbound requests in flight
timeout_ms     = 20000   # per HTTP attempt

[[models]]
name              = "openai-text-embedding-3-small"  # route name
provider_model_id = "text-embedding-3-small"         # API id
dim               = 1536   # responses are checked against this
max_batch         = 512    # host sub-batches to this
max_tokens        = 8191   # advertised as sequence_len
# space    = "openai-text-embedding-3-small"  # vector space; default is name
# priority = 1                                # lower wins; omit for derived order
# added    = "2026-09-08T00:00:00Z"           # RFC 3339; unstamped routes precede stamped ones

base_url replaces the origin only. Each connector appends its own path. OpenAI-shaped APIs and UniVec embeds use /v1/embeddings. UniVec conversion uses /v1/convert, Cohere uses /v2/embed and Gemini uses its batch path. The URL must be absolute http or https with a host. It must not carry a query string, fragment, userinfo or trailing /.

A reverse proxy, a gateway, a self-hosted OpenAI-compatible endpoint or Azure's v1 API under /openai/v1 all work as a base_url. Azure OpenAI's classic API needs a deployment path and an api-version query, which this schema has no field for.

Plain http to anything but loopback is refused unless the file opts in:

toml
base_url = "http://vllm.internal:8000"
allow_insecure_transport = true   # required, and never the default

A file that opts in still warns on every doctor run. Loopback needs no opt-in.

Referenced secret paths must be absolute. The CLI, the embedded launcher and postvec-server all resolve them and none share a working directory.

Key sources ​

Exactly one of three, mutually exclusive:

FieldUse it whenNote
api_key_file = "/path"The usual choiceAn absolute path to a regular file (not a symlink, not hard-linked) at 0600. A trailing newline is trimmed
api_key_env = "OPENAI_API_KEY"Containers, systemd unitsResolved in the inference process's environment. The postmaster's or the unit's, not your shell's
api_key = "sk-..."Last resortThe value sits in the 0600 file

api_key_file is opened with O_NOFOLLOW and the mode is checked on the descriptor the host actually opens. Kubernetes projected secret volumes are symlinks into a ..data directory and are mounted world-readable, so they cannot be referenced this way. Use api_key_env there.

AWS Bedrock ​

AWS replaces api_key* with region plus either a Bedrock bearer token (bearer_token, bearer_token_file, bearer_token_env) or a static SigV4 pair (access_key_id and secret_access_key, each with its own _file / _env variant):

toml
provider = "aws"
region   = "us-east-1"
bearer_token_file = "/etc/postvec/keys/bedrock.key"

[[models]]
name              = "aws-titan-embed-text-v2-0"
provider_model_id = "amazon.titan-embed-text-v2:0"
dim               = 1024
max_batch         = 1        # Titan invokes one text per request

The signer takes static credentials only: no session tokens, no instance profile, no IMDS.

A few AWS specifics:

  • The connector speaks the Amazon Titan embedding schema. Other vendors hosted on Bedrock use different body shapes and are not in the built-in catalogue.
  • region is the whole endpoint. The connector builds bedrock-runtime.<region>.amazonaws.com from it, so base_url does not apply to aws and provider add refuses the flag. A region is restricted to [a-z0-9-].
  • Bedrock's InvokeModel sends one text per request, so max_batch is
    1. On a Titan-backed column, lower postvec.batch_size rather than raising the timeout.
  • provider add and provider test probe the bearer-token variant. For a SigV4 file, pass --no-verify and confirm with provider ls plus a first write.

Gemini and Cohere dimensions ​

Where a provider documents it, dim is also requested: Gemini receives outputDimensionality and Cohere v4 output_dimension.

  • Cohere v4 produces 256, 512, 1024 or 1536 and nothing else. A descriptor asking for another width is refused at load.
  • Cohere v3 widths are fixed per model (embed-english-v3.0 and embed-multilingual-v3.0 are 1024, the -light- pair 384) and are checked exactly.
  • Gemini ships one model: gemini-embedding-001, with dim anywhere in Google's documented 128..=3072. Any other Gemini model id is refused at load. A reduced Gemini vector is not normalised by Google, so postvec renormalises it.

Providers that separate query text from stored text send search_query for search() and search_document for worker writes, embed() and migration re-embeds. Cohere, Gemini and UniVec use this distinction.

UniVec converter entries ​

UniVec is the only connector that accepts kind = "convert". An embed entry uses the common schema above. A converter adds the source and target names in the provider's vocabulary and in postvec's vocabulary:

toml
provider = "univec"
api_key_file = "/etc/postvec/keys/univec.key"

[[models]]
name               = "univec-convert-snowflake-to-bge-m3"
kind               = "convert"
provider_source_id = "snowflake-arctic-embed-l-v2.0"
provider_model_id  = "baai-bge-m3"
source_model       = "snowflake-arctic-embed-l-v2.0"
target_model       = "baai-bge-m3"
source_dim         = 1024
dim                = 1024
max_batch          = 96

provider_source_id and provider_model_id go on the UniVec request. source_model and target_model are matched against a postvec column and the requested migration target. source_dim validates every input vector; dim validates every output vector.

Converter entries refuse max_tokens. They also refuse a missing name or dimension, identical source and target names, or a connector type other than univec. One invalid entry prevents the whole file from loading.

Hosted converters are direct routes for migrate() and convert(). Embed-bridge resolution uses local models. To search a retired space without migrating, see search a retired space. UniVec hosted models has the hosted conversion sequence.

Loading rules ​

  • *.toml in the directory, read in lexicographic order. Dot-files skipped.
  • A file that fails to parse, or whose key cannot be resolved, is skipped with a logged error. Every other provider and every local model keeps working.
  • A file is checked as a whole. An unknown provider type, a connector with no credential, an aws file with no region, an implausible dim or an unusable base_url skips the entire file, including its other models.
  • The schema is per connector. Unknown fields for that connector type are refused.
  • enabled = false parses and serves nothing.
  • A public name may appear once. Two entries in one file: the file is refused. Two files claiming the same name: neither serves until one drops it.
  • A name that a loaded local engine model already serves refuses the whole file at load. The local model keeps serving under that name.
  • The directory is bounded as a whole: at most 32 connector files, 512 provider models and 256 total max_concurrent across every file. A single file holds at most 256 [[models]] entries (one whole UniVec catalogue with headroom), is at most 256 KiB, and a referenced secret at most 16 KiB. Breaking a directory-wide ceiling fails the whole scan.

provider add and provider rm validate the directory as it would be after the write. A file the host would skip as a whole is refused before it is written, so existing models on that provider stay up.

The directory ​

postvec setup --embedded creates /etc/postvec/providers.d (0700, owned by the cluster owner), and so does the first provider add. The CLI creates it, because only the CLI knows which account owns the cluster. An absent directory is the default.

On a postvec-server node the directory is <server-root>/providers.d, and a key copied from postvec login lands in <server-root>/keys/<stem>.key. Run postvec provider add --path <server-root> on the node to write the file there.

The directory is owned by the account that reads it (or by root) and must not be group- or world-writable. The same account is the only one besides root that can rewrite the directory or any ancestor. Anyone who can write there can drop in a connector file and choose where this host sends source text or stored vectors. The serving host refuses those cases. provider add also refuses to write there, and doctor fails provider.directory. Read and execute bits only disclose which providers are configured; those stay a warning.

provider add and provider rm take an advisory lock on the directory for the whole read-modify-write, so two administrators running them at once cannot lose one another's change.

Change the location with postvec.providers_path (SIGHUP) or setup --embedded --providers-path DIR, and on a postvec-server node with --providers-path or POSTVEC_PROVIDERS_PATH.

postvec uninstall reports connector files and leaves them. Package removal does the same.

Model names ​

The built-in catalogue maps each --model id to a public route name and a vector space. same means the space equals the route name.

ProviderModel idRoute nameSpaceDim
openaitext-embedding-3-smallopenai-text-embedding-3-smallsame1536
openaitext-embedding-3-largeopenai-text-embedding-3-largesame3072
openaitext-embedding-ada-002openai-text-embedding-ada-002same1536
googlegemini-embedding-001gemini-embedding-001same3072
cohereembed-v4.0cohere-embed-v4-0same1536
cohereembed-english-v3.0cohere-embed-english-v3-0same1024
cohereembed-multilingual-v3.0cohere-embed-multilingual-v3-0same1024
awsamazon.titan-embed-text-v2:0aws-titan-embed-text-v2-0same1024
awsamazon.titan-embed-text-v1aws-titan-embed-text-v1same1536
mistralmistral-embedmistral-mistral-embedsame1024
openrouteropenai/text-embedding-3-smallopenrouter-openai-text-embedding-3-smallopenai-text-embedding-3-small1536
openrouteropenai/text-embedding-3-largeopenrouter-openai-text-embedding-3-largeopenai-text-embedding-3-large3072
openrouteropenai/text-embedding-ada-002openrouter-openai-text-embedding-ada-002openai-text-embedding-ada-0021536
openroutergoogle/gemini-embedding-001openrouter-google-gemini-embedding-001gemini-embedding-0013072
openroutercohere/embed-v4.0openrouter-cohere-embed-v4-0cohere-embed-v4-01536
openroutercohere/embed-english-v3.0openrouter-cohere-embed-english-v3-0cohere-embed-english-v3-01024
openroutercohere/embed-multilingual-v3.0openrouter-cohere-embed-multilingual-v3-0cohere-embed-multilingual-v3-01024
openroutermistralai/mistral-embedopenrouter-mistralai-mistral-embedmistral-mistral-embed1024

These 18 entries are the whole catalogue. An OpenRouter entry mirrors a vendor route, so it shares that route's space: a column bound to openai-text-embedding-3-small is served by openrouter-openai-text-embedding-3-small once only the OpenRouter key exists.

Any other id works too. The probe measures the dimension, or --dim states it. A space has one width: two embed entries in one file that disagree refuse the file; an entry that disagrees with another file or with a loaded local model is skipped at load time (the rest of its file serves), and the cache refresh skips a contradicting row with a WARNING.

The built-in catalogue covers the six connectors above, not UniVec. provider add univec --model baai-bge-m3 takes the dimension from UniVec's catalogue, or measures it for an id the catalogue lacks, and writes route univec-baai-bge-m3.

For an unlisted id, the name is derived mechanically: lowercase, and every character outside [a-z0-9._-] becomes -. provider_model_id keeps the id exactly as the API expects it.

gemini is accepted as an alias for google, and amazon for aws. The file always records the canonical name. An alias never changes what you type in SQL.

Commands ​

CommandResultNetwork
provider add TYPE --model ID...Write or extend the file, verify, reload the host, refresh postvec.modelsOne embed per verified model, unless --no-verify
provider add univec --convert-to MODELWrite or extend the file with every catalogue converter into MODEL, verify, reload and refreshOne vector conversion per added entry, unless --no-verify
provider lsProviders, key sources, models, dims and whether the host serves them nowLoopback
provider test NAME [--model ID]Verify embed or converter entries on demandOne embed or vector conversion per selected entry
provider rm NAME [--model ID]Drop one model entry or the whole file, then reloadLoopback

Useful options on add:

OptionEffect
--name STEMWrite STEM.toml instead of the type's name. Two OpenAI-compatible endpoints, two files
--api-key-file P / --api-key-env VAR / --key-stdinChoose the key source. Without any of them and with a TTY, a hidden prompt asks and records the key inline as api_key
--api-key-from-loginunivec only. Copy the key postvec login stored into <providers root>/keys/<stem>.key and use it for inference. An interactive run offers this; a script must ask. Refused when the file holds a different key
--replace-copied-keyWith --api-key-from-login: rotate a key file this command copied for this connector earlier. The new key is verified before the old one is replaced
--base-url URLGateways, Azure-shaped fronts or a mock server. Not accepted for aws
--regionRequired for aws, and only valid there
--dim NVector dimension for a single model, and the requested width where the provider takes one (Gemini, Cohere v4). Required with --no-verify when the catalogue and the probe cannot supply it
--space SPACEVector space this route serves. Default: the catalogue space, else the public name
--preferGive the new route priority 1 in its space, ahead of the local model
--convert SRC:DSTunivec catalogue converter, by provider id pair. Repeatable
--convert-to MODELunivec: every catalogue converter into MODEL. Repeatable
--convert-from MODELunivec: every catalogue converter out of MODEL. Repeatable
--all-convertersunivec: every catalogue converter
--no-catalogunivec: skip the catalogue fetch. Every dimension then comes from --dim or the probe
--path DIRWrite into DIR/providers.d instead of the selected cluster. Default: POSTVEC_PROVIDERS_PATH, then the cluster's providers path
--acknowledge-in-useAccept that columns bound to these names start sending their text to the provider at the next worker cycle. Required with --yes when a column is affected
--yesSkip the confirmation prompt. Required for a mutation without a TTY
--no-verifySkip the live probe
--dry-runPrint the plan and change nothing
--convert-source ID --convert-target IDHidden. Manual converter: UniVec provider ids
--source-model NAME --target-model NAME --source-dim NHidden. Manual converter: postvec route names and the input dimension
--converter-name NAMEOverride the derived univec-convert-<source>-to-<target> name

The plan is confirmed before the probe. Declining it, or failing the in-use acknowledgement, makes no API call.

add and test use an embedding probe for embed entries and a single-vector conversion probe for converter entries. A connector change, such as a new key source or base_url, rechecks every entry in that file.

Target resolution:

TargetBehaviour
--path DIRFilesystem management of DIR/providers.d, or of DIR itself when it already is one. DIR must already exist. New files inherit its owner. An embed entry requires --acknowledge-in-use; a new converter skips that flag
Embedded clusterManage postvec.providers_path, scan the databases for affected columns, reload the running host
Remote fleet (grpc)Refused by name on the database host, because connector files live on the inference nodes. The message points at --path; run provider add --path <server-root> on a postvec-server node

Reading provider ls:

text
providers.d: /etc/postvec/providers.d
openai  (openai, key: file:/etc/postvec/keys/openai.key)
  openai-text-embedding-3-small                dim 1536   served
cohere  (cohere, key: env:COHERE_API_KEY)
  cohere-embed-v4-0                            dim 1536   NOT served (reload or restart the host)
azure  (openai, key: env:AZURE_KEY)  [enabled = false]
  openai-text-embedding-3-large                dim 3072   disabled in the file
mistral  (mistral, key: none)
  ! the host REFUSES this file: provider "mistral" needs an API key: set
    exactly one of api_key_file (recommended), api_key_env or api_key
  mistral-mistral-embed                        dim 1024   not loadable (see above)

NOT served after a hand-edit: reload or restart. enabled = false: reported as disabled. REFUSES: the file will not load. ls and doctor apply the loader's rules.

doctor checks ​

postvec doctor has a provider.* family, all read-only:

CheckReports
provider.directoryFails when the directory is group- or world-writable, or when ownership or ancestors fail the write-safety rules. Warns on mere read/execute bits. "Does not exist" is a pass
provider.fileA file the serving host would refuse: mode, unknown fields, two sources for one secret, an unknown type, no credential, a bad dim / region / base_url, no [[models]]. Warns on a plaintext base_url to a non-loopback host
provider.key-sourceA referenced key file that is missing, a symlink or too permissive, or a named variable that is absent. Every source, including both halves of an AWS SigV4 pair
provider.servedA configured model missing from the running host. Skipped for enabled = false files

A complaint about an environment variable can be a false alarm. doctor observes its own environment, and the postmaster's is what matters. The check says so.

provider add / rm write the files first and then ask the host to reload. When no host answers on a cluster target, the result is partial (exit 3): the files are correct and a restart applies them. On a --path target the same situation is a note, not a partial result.

Limits ​

TopicWhat the product does
CredentialsKeys live in providers.d on the inference host. PostgreSQL holds a path
Spend controlmax_concurrent bounds calls in flight. Set quotas on the provider
AWS authStatic SigV4 pair or a Bedrock bearer token
Azure OpenAIbase_url fronts the /openai/v1 API
Geminigemini-embedding-001 only (the documented contract)
Hosted convertersDirect routes for migrate() and convert(). Embed-bridge uses local models
Provider TLSrustls with the bundled Mozilla root set
Inference transportLoopback gRPC (embedded) and postvec-server gRPC (remote) are plaintext. Restrict the gRPC port. Use provider-side quotas as the spend control