AI research assistant · RAG SaaS
Retrivo Vault
The research assistant that only knows what you’ve read.
- Role
- Solo · design, build and ship
- Focus
- AI · Web
- 142
- tests
- 3-tier
- chunking cascade
- 15s
- token-replay grace
- 6
- input formats
Overview
A production-grade, individual-focused RAG SaaS. Upload contracts, papers and notes, ask questions in plain language, and get streaming answers that cite the exact passage they came from — and say “not in your documents” instead of guessing.
What it does
- PDF, TXT, Markdown, DOCX, CSV and web pages by URL, ingested into a private vault
- Streaming, citation-backed chat with match scores and clickable inline sources
- Free / Pro / Max plans with trial-aware quotas and Stripe billing
- In-app, email and Web Push notifications; activity log and data export
- Personal API keys, HMAC-signed webhooks and an OpenAPI spec
- Operator admin console: presence, complimentary grants, broadcasts, audit log
Engineering decisions
Isolation is enforced by the index, not the query
The userId filter is part of the Atlas Vector Search index definition and applied inside the ANN search — post-filtering would let other users’ vectors consume the top-k slots.
Chunking that survives hostile documents
Paragraphs → sentences → fixed-width hard split, with 150-character overlap so a fact straddling a boundary appears whole in at least one chunk.
Refresh-token reuse detection with a grace window
Token families with rotation stamps: a replay within 15s is a client retry, anything older revokes the entire family.
Stack
- Backend
- NodeExpressMongoDB Atlas Vector Search
- AI
- Gemini embeddingsGemini generationSSE streaming
- Frontend
- ReactTypeScript
- Platform
- StripeWeb PushOpenAPI