Permission-aware RAG: filter before retrieval, or after?
A retrieval system that serves more than one person has to decide what each person may see. It can make that decision before the search ranks anything, or after. The choice decides what can leak and where, and it is the first thing to ask of any product that says it does permission-aware RAG. This page defines both, shows where products document their choice, and ends with a checklist for buyers.
Updated
01Before and after, defined
Pinecone's guide to RAG access control (January 2026) puts the two approaches side by side.
- Before, or pre-filtering: "Only the list of documents that the user can access is embedded and sent to the vector database." The search never sees what the person may not read.
- After, or post-filtering: "A CheckPermissionRequest is performed on every document ID that the vector database returns." The search ranks everything, and the results are trimmed.
Pinecone's own advice on choosing: "If you have a high positive hit-rate of documents from your vector database, a post-filter approach works well. Conversely, if you have a large corpus of documents in your RAG pipeline and a low positive hit-rate, the pre-filter approach is more efficient."
Filtering before retrieval has a known trap with approximate vector indexes. The pgvector documentation says: "With approximate indexes, filtering is applied after the index is scanned. If a condition matches 10% of rows, with HNSW and the default hnsw.ef_search of 40, only 4 rows will match on average." A person who may read a tenth of the documents asks for 40 results and gets about 4. pgvector 0.8.0 added iterative index scans, which "automatically scan more of the index until enough results are found". Any product that filters inside an approximate index should say how it handles this.
A permission rule can also live in the database itself. Supabase's guide pairs pgvector with PostgreSQL row-level security, so "you can restrict which documents are returned during a vector similarity search to users that have access to them".
02What products document
Where each product's own pages place the permission check, as of September 2026:
| Product | When permissions apply | In its words |
|---|---|---|
| Amazon Bedrock Knowledge Bases with Verified Permissions | Before | "Amazon Bedrock Knowledge Bases applies the metadata filter before the vector similarity search runs." |
| Nasuni AI Activate | Before | "Each request is scoped to what the signed-in user is already allowed to see, before retrieval, not after." In invite-only preview, with general availability targeted for the fourth quarter of 2026. |
| Cerbos | Before, compiled into the search | "Cerbos query plans translate into native filter syntax for Pinecone, Weaviate, Chroma, Qdrant, and FAISS." |
| SpiceDB | Either | Its guide shows how to "pre-filter and post-filter vector database queries with a list of authorized object IDs". |
| OpenFGA | After | "OpenFGA filters the candidates: for each candidate doc:X, check whether user:Y has can_view." |
| Archestra | After ranking | Its documented steps expand the query, search, fuse and rerank, and only then "Filter by access. Chunks the asking user cannot read are removed." |
| Onyx | At query time | "Source permissions enforced at query time." The pages read do not say whether that is before or after ranking. |
| PipesHub | At query time | "Access is resolved when the query runs, against the source system's own permissions, instead of being approximated at build time." |
| Elasticsearch connectors | Stored with each document | The network drive connector keeps each file's access list once document level security is enabled: "Permissions are not synced by default." |
Two different questions hide under "when". One is when the permissions are read: copied into the index at sync time, or asked of the source system live. The other is when they are applied: before the ranking or after it. A product can read permissions live and still apply them after ranking. Ask both.
03What filtering after ranking costs
- Fewer results than asked. If the top 20 are ranked across everything and then trimmed, a person with narrow access can be left with two, or none, while better answers they may read sat at position 25.
- Restricted text passes through other models first. In a pipeline that reranks before it filters, the chunks a person may not read are sent to the reranking model before they are removed. If that model is hosted, restricted text leaves the building on every question.
- Signals leak. Result counts, timing and "no results" answers can hint that something the person may not read exists.
Filtering before ranking avoids all three, at the price of doing the permission work on every query, and of handling the approximate-index trap above.
04Beyond who may read
Permission is one gate. A retrieval system that answers from a company's files can hold back more than what a person may not read:
- Superseded material. An old version of a policy is readable, but it should not be served as the current answer. Microsoft's SharePoint Advanced Management states that "Copilot isn't trained on archived content", which applies to whole archived sites.
- Unreviewed material. Guru says its MCP tools "do not access raw documents or unverified data, they rely on Guru's cited, permission-aware knowledge layer."
- Stale material. Bloomfire "uses AI to flag outdated or redundant content before it pollutes your search results", and triggers a review.
05A buyer's checklist
- Is the permission check applied before the search ranks anything, or after?
- Are permissions read from the source system, or written into the product by an administrator? How fast does a removed permission take effect: the next question, or the next sync?
- Does any document text reach a model (a reranker, a query rewriter, an embedding service) before the permission check? Where does that model run?
- If the search uses an approximate vector index, what happens when a person may read only a small share of the documents?
- Is access enforced a second time, in the database, in case the application has a bug?
- Are superseded, expired or unreviewed documents kept out of answers by default?
- Can you write a test that names a person and the files that must never reach them, and does it fail loudly when one does?
- What does the log keep for each answer: who asked, what they asked, what they were given, and where the model that received it runs?
Where PremAgentic fits
PremAgentic filters before retrieval. It indexes the Markdown, text, PDF, Word and Excel files your organization keeps and applies four gates as hard SQL conditions before anything is ranked, so no relevance signal can outweigh them:
- Access. Set in PremAgentic, folder by folder: ordered allow and deny entries, where the first entry that names the caller decides and no entry means no. Checked on every search.
- Lifecycle. Superseded material is held back unless a caller asks for history.
- Trust. A document that declares a machine author is held from agents until a person has reviewed it.
- Freshness. Past its stale-after date, a document is flagged for people and held from agents.
In an install made with its setup command, PostgreSQL row-level security enforces access a second time. A deployment's golden set can name a person and the files a question must not reach, and the check fails loudly if one does. Every question is logged with who asked, the text of the question, where the assistant's model ran and the passages returned.
It runs on your own servers, Windows or Linux, on stock PostgreSQL 14 or later with no extension, and a local embedding model. PremAgentic is open source under the GNU Affero General Public License 3.0. The code is at github.com/premagentic/premagentic, and the latest release has the Linux and Windows archives.
Sources
Every fact about another product comes from that product's own pages, read on September 27 and 28, 2026.
- Pinecone, RAG with access control (2026-01-08)
- pgvector, filtering and iterative index scans
- Supabase, RAG with permissions
- AWS, secure multi-tenant RAG with Amazon Bedrock and Verified Permissions
- Nasuni AI Activate; Nasuni press release (2026-04-07)
- Cerbos, access control for RAG
- SpiceDB, secure RAG pipelines
- OpenFGA, RAG authorization
- Archestra knowledge base
- Onyx
- PipesHub README
- Elastic network drive connector
- Microsoft, SharePoint Advanced Management for Copilot
- Guru MCP server
- Bloomfire