---
topic: system-design
author: Crashtech Editorial
date: Oct 7, 2026 · read: 2 min
updated: October 8, 2026
---

pgvector and Hybrid Search: When Keeping Retrieval in Postgres Helps

Postgres can combine vector and lexical retrieval with rank fusion, simplifying consistency without replacing every dedicated search system.

– –

Dual-database split architecture versus unified PostgreSQL hybrid search

Conceptual architecture; arrows show relationships, not measured latency, energy or safety guarantees.

Consistency has a boundary

A separate search store creates a derived-data pipeline. Change-data capture and outboxes can make that pipeline reliable, but lag, deletion and recovery need explicit contracts. Divergence is a risk to manage, not an inevitable permanent failure.

Putting a document and its vector in one PostgreSQL transaction can remove that particular cross-store write boundary. It does not make an asynchronously generated embedding instantly current. Track content versions so an old embedding job cannot overwrite a newer document’s representation.

Use two retrieval signals

Dense similarity helps with paraphrases. Lexical matching helps with words, product codes and names. Exact identifiers may need explicit equality or structured filters: stemming full-text search does not guarantee exact identifier matching.

PostgreSQL provides tsvector, tsquery and ranking functions including ts_rank_cd. A GIN index can accelerate matching, but the ranking function is not BM25. pgvector supplies vector operators and approximate indexes such as HNSW. [1] [2]

Fuse bounded candidate lists

Reciprocal Rank Fusion combines positions rather than requiring scores from different systems to share a scale. The contribution of a result at rank r is 1/(k+r). It is a useful baseline; relevance evaluation should decide whether it beats calibrated scoring or a reranker. [3]

Illustrative retrieval pipeline:
1. Authorize the request and scope both searches.
2. Obtain a bounded vector candidate list in distance order.
3. Obtain a bounded lexical list in relevance order.
4. Assign ranks within those lists.
5. Sum 1 / (60 + rank) for each candidate's appearances.
6. Recheck visibility and return the highest-scoring results.

The constant 60 is an example, not an optimum for every corpus. Limit candidate retrieval before applying window functions, and inspect the query plan: an outer LIMIT alone does not guarantee an efficient nearest-neighbor index scan.

Access control and recall need separate tests

Apply authoritative authorization to both branches and returned records. Row-level security can help when the application role and policies are configured correctly; privileged and owner roles need particular care. [4]

With approximate vector indexes, filtering can leave fewer candidates than requested. pgvector documents iterative scans and tuning options, so test restrictive tenant filters rather than assuming all filtering occurs before ANN retrieval. [1]

Measure recall, latency, update rate and operational load. Choose a dedicated engine when its capabilities justify the extra system. There is no universal five-millisecond guarantee or 500-million-vector threshold that decides the architecture for every team.

Advertisement

Frequently asked questions

Is PostgreSQL ts_rank_cd a BM25 implementation?

No. PostgreSQL’s built-in ts_rank and ts_rank_cd use their documented ranking functions. BM25 requires a suitable additional implementation or search system.

Does a tenant filter guarantee high vector recall?

No. Authorization and retrieval quality are separate. Approximate index scans can return too few filtered candidates; evaluate recall, candidate budgets and iterative scans for the actual query.

Sources & further reading

/* Comments */