Vector search is an indexing decision, not a database purchase
I still think vector search is an indexing decision, not a database purchase—but “choose HNSW or IVF” is too small a reading of that claim. The production decision includes the query contract, tenant boundary, embedding identity, recall evidence, rebuild path, and operating owner around the index. A specialist service may package those responsibilities well. It does not make them disappear.
The algorithm remains one part of the contract. HNSW buys strong recall and low latency with a memory-heavy graph; IVF and product quantization trade more approximation for a smaller footprint. Those choices should follow a measured workload: corpus size, update rate, latency budget, memory ceiling, and the acceptable probability of missing a relevant item. Approximate-nearest-neighbour recall is a product SLO, not a one-time benchmark score. Measure it against labelled query classes and keep measuring after data, filters, and embeddings change.
Filtering is where the index becomes application architecture. “Nearest neighbours where tenant_id = X, region = Y, and effective_at covers today” can behave very differently under pre-filtering, post-filtering, or in-graph filtering. A post-filter may satisfy isolation yet quietly destroy recall for a sparse tenant. An unguarded pool may retrieve across tenants. I want mandatory scope injected below the caller, filter selectivity included in evaluation, and failure closed when the isolation predicate cannot be proved. Vector tenancy fails at metadata isolation before it fails at similarity.
Embedding identity belongs in the index schema. Model, dimensionality, preprocessing, task prefix, and source version determine what a stored vector means. Switching any of them is a data migration because old and new representations are not safely comparable merely because both are arrays of numbers. A production index should reject query/index mismatch and support dual-read cutovers: build the candidate representation beside the current one, compare retrieval outcomes, promote deliberately, and preserve rollback until the new path earns trust.
The database choice follows from those obligations. Keeping vectors beside relational or document data can remove a synchronization boundary, reuse access control, and make hybrid queries easier. A dedicated engine can earn its place when scale, latency, isolation, or operational tooling exceeds what the existing database can supply. Either way, derived vectors need lineage to source records, deletion propagation, backfill checkpoints, index-health evidence, and an owner for compaction and rebuilds. “Managed” transfers some operations; it does not transfer semantic accountability.
Portability is part of the lifecycle too. Index files may be engine-specific and rebuildable, but the durable inputs should not be trapped: source identifiers, canonical text or features, metadata, embedding-version records, and evaluation queries need exportable forms. Exportable vector data keeps an engine reversible and makes disaster recovery more than a promise to re-embed everything from memory.
There is one precise concession: when vector retrieval is itself the product, operates at very large scale, and has strict latency and recall targets, choosing a specialist distributed vector system can be the correct database purchase. The boundary is still operationally testable—the specialist must outperform simpler colocated indexes on the real filtered workload enough to justify another security model, failure domain, and synchronization path.
I now evaluate vector search as an index lifecycle. Define query classes and isolation first; bind every vector to its representation version; measure filtered recall and latency; rehearse rebuild, migration, export, and deletion; then choose the engine. Retrieval quality is mostly decided above and around the index—by the pipeline feeding it and the contracts governing change. Database branding is the least durable part of that design.