Vector lock-in and embedding debt
Stored vectors are bound to the model that produced them.
Vector lock-in
An embedding is a point in the vector space of the model that produced it. openai-text-embedding-ada-002, BGE-M3, Gemini, Arctic and GTE each live in a different space, even when dimensions match. Embed a query with model B and search a corpus from model A, and the ranks are invalid.
Once the stored corpus, its ANN index and every query caller assume space A, the model is part of the data contract. That binding is vector lock-in.
Embedding debt
Changing that contract normally means a corpus re-embed, an index rebuild, dual writes during cutover, a fresh retrieval evaluation and a rollback plan. Deprecation of the source provider or model makes the same work time-sensitive. Every new vector written in the old space grows the eventual migration.
That accumulated, deferred migration work is embedding debt.
- Search a retired space converts each query vector into the stored space (embed-bridge). The corpus stays.
search()plus filters and BM25 combine lexical and semantic retrieval.- Embedded mode runs local open-weight models on the database host.
- postvec-server runs the same models in a separate process, on a CPU or GPU fleet or on managed PostgreSQL. Source text and query strings travel to those nodes.
If later you want the stored contract itself to change, migrate() converts stored vectors. The column's model contract then changes.