The short version
Put the user filter inside the vector search itself. If you search everything and remove other people’s chunks afterwards, those chunks still use up the slots you asked for, and one forgotten line of code puts a stranger’s document into someone else’s answer.
The version most apps ship first
The obvious way to build a multi user RAG app is to search the whole collection, take the best matches, and throw away anything that does not belong to the current user. It passes every test you write on your own laptop, because on your laptop there is only you.
Now do the arithmetic with real users. Fifty people have uploaded contracts, and every one of those contracts has a termination clause. You ask for the 8 nearest chunks to “how do I terminate early?”. Those 8 chunks are scattered across many users, because that is what “nearest” means. After the filter you keep one or two, or none, and the assistant answers from almost nothing. Nothing crashed, so nothing alerts you. Answers just get worse as you grow.
There is a second problem. The only wall between your users is a filter that lives in application code, and application code gets refactored by tired people.
Put the filter inside the search
In Retrivo Vault the userId filter is part of the Atlas Vector Search index definition and runs inside the nearest neighbour search. Vectors that belong to someone else are never candidates, so the search returns the 8 best chunks from your documents because those are the only chunks it looks at.
It takes two pieces. The index declares userId as a filter field next to the vector field:
{
"fields": [
{ "type": "vector", "path": "embedding", "numDimensions": 768, "similarity": "cosine" },
{ "type": "filter", "path": "userId" }
]
}And the query passes the filter to the same stage that does the searching. numDimensions has to match your embedding model, so 768 here is only an example.
export async function searchChunks({ userId, queryVector, limit = 8 }) {
if (!userId) throw new Error('searchChunks needs a userId')
return chunks
.aggregate([
{
$vectorSearch: {
index: 'chunks',
path: 'embedding',
queryVector,
numCandidates: limit * 25,
limit,
filter: { userId },
},
},
{ $project: { text: 1, source: 1, score: { $meta: 'vectorSearchScore' } } },
])
.toArray()
}Two details matter. The function refuses to run without a userId, so a bug upstream fails loudly instead of searching everyone. And numCandidates is how many nearest neighbours the engine considers before returning the top limit. A larger number improves recall and costs latency, so tune it with your own data rather than copying mine.
How to prove isolation works
Do not trust this because a blog post said so. Write the test that would catch the failure, and keep it in CI.
- Create two users and upload the same document to both. Identical text gives identical vectors, which is the hardest case for isolation.
- Search as user A and assert that every returned chunk has A’s userId. Assert it for B as well.
- Give user B one tiny document and search with limit 8. You should get a few results, not eight and not zero. This catches post filtering, which returns fewer results for small accounts.
- Call the search function with no userId and assert that it throws.
Chunking that survives messy uploads
People upload scanned contracts, CSV files with one enormous line, and pages with no paragraph breaks. A splitter that only cuts on blank lines hands you one giant chunk and useless retrieval, so Retrivo Vault splits in three tiers: paragraphs first, then sentences, then a fixed width cut as the last resort. Chunks overlap by 150 characters, so a fact that sits on a boundary still appears whole in at least one chunk.
Here is a sketch of the idea. The real thing also packs small pieces together up to the size limit, but the fallback order is the part worth copying.
const MAX = 1000
const OVERLAP = 150
export function split(text) {
const out = []
for (const paragraph of text.split(/\n{2,}/)) {
if (paragraph.length <= MAX) { out.push(paragraph); continue }
for (const sentence of paragraph.split(/(?<=[.!?])\s+/)) {
if (sentence.length <= MAX) { out.push(sentence); continue }
for (let i = 0; i < sentence.length; i += MAX - OVERLAP) {
out.push(sentence.slice(i, i + MAX))
}
}
}
return out.filter((piece) => piece.trim())
}Let the assistant say it does not know
Answers come only from retrieved passages. Every citation shows the exact passage and its match score, and when the answer is not in your documents the assistant says exactly that. A user who has seen it admit ignorance once will believe it the next time it sounds sure. A confident wrong answer costs more trust than ten honest ones earn.
The same idea shows up in tokens
Refresh tokens follow the same rule. They are issued in families and each rotation is stamped. If an old token turns up again within 15 seconds it is treated as a client retry. Anything older revokes the whole family, because by then someone else probably has a copy. Clients misbehave, so the safe outcome has to be the default one.
The rest of the engineering notes, and the 142 tests behind them, are in the Retrivo Vault case study.