RAG development, by people who have run one in production
Most RAG projects fail on retrieval quality, permissions and latency, not on the model. A retriever that returns the right chunks to a user who was never allowed to see them is a breach, not a feature. A pipeline that takes twelve seconds to start answering is a pipeline nobody uses. We have built the kind that holds up: authorization inside retrieval rather than after it, a first token budget worked backwards from a target, and per layer logging so a wrong answer can be traced to the layer that caused it.
Who this is for
- - A SaaS product with a large document base, customer support load, and an assistant that is slow, wrong, or both.
- - An enterprise with permissioned documents, where the assistant must never show a user something their role cannot see.
- - A team that built a RAG demo in a notebook and cannot get it to behave on real users and real documents.
- - A company choosing between RAG, fine tuning and long context and wanting a straight answer rather than a vendor’s.
What you get
- - An eight layer architecture with a written contract between layers: ingest, index, query understanding, authorized retrieval, ranking, generation, verification, observability.
- - Permissions enforced inside retrieval: the user’s identity filters candidates before ranking, not after the answer.
- - A first token latency budget per layer, measured at the median and the slowest five percent.
- - An evaluation set of real questions with graded answers, run before every change.
- - The pipeline, the index configuration and the evaluation set in your accounts.
What it costs
There is no honest fixed price before discovery, and anyone who gives you one is guessing. What we can say in writing after two weeks: the build cost, the running cost per month at your expected volume, and what moves each. The build is driven by how many systems the work must touch and how strict the accuracy, latency and compliance bars are. The running cost is model calls plus infrastructure, and it can differ by ten times depending on the choices made in week one. How we price AI work.
How long it takes
Discovery, two weeks: we map the problem, the systems involved and the success criteria, and put the numbers in writing. Pilot, four to six weeks: one capability, live, on a slice of real traffic, measured against those criteria. Then production hardening and scale out. Each stage is a separate agreement with no minimum commitment, so you can stop after discovery with a plan you own. How we work.
What you own afterwards
Everything. The code, the prompts, the evaluation set, the infrastructure configuration and the documentation, in your accounts, under your name. We do not hold anything hostage and we do not license our own platform to you. The evaluation set matters most: it is the thing that lets your team, or any other team, change the system later without breaking it. The deliverables clause.
The one we point to
A US automotive SaaS assistant missed intent and took 10 to 30 seconds to begin answering over 120 GB of documents, roughly 40,000 PDFs. We rebuilt it: intent detection and routing before retrieval, authorization aware retrieval, caching and targeted context construction, on Google Cloud with Gemini. First token in under three seconds in typical interactions, about 98 percent accuracy across intent and retrieval in the client’s evaluated use cases, and operating cost down by roughly 30 percent.
Read the caseQuestions buyers ask
What should a RAG development company be able to show you?
A production system, not a notebook: how permissions are enforced at retrieval time, the first token latency at the median and the slowest five percent, how a wrong answer is traced to the layer that caused it, and an evaluation set. If a vendor cannot answer those four, they have built demos.
How long does a RAG system take to build?
Discovery takes two weeks: we read a sample of your documents, map the permission model and the question types, and agree the latency and accuracy targets. A pilot takes four to six weeks on one document set with real users. Production hardening and scale out follow. Each stage is a separate agreement.
How much does RAG development cost?
The build is driven by document volume and variety, the complexity of the permission model, and the latency target. The running cost is embedding, retrieval and generation per query, and caching usually cuts it sharply. Both go in writing after discovery; on the system above, operating cost fell by roughly 30 percent after the rebuild.
Do we still need RAG with large context windows?
Usually yes. Long context changed where the line sits, not whether it exists: stuffing the window is now right for small, static, unpermissioned corpora. Permissions, freshness, cost per query and traceability still make retrieval necessary for most business document bases.
Before you decide
If you want the technical background first, read how retrieval augmented generation works. Then, from our engineers:
Tell us about your documents and who may see them
Two weeks of discovery puts the cost, the timeline and the plan in writing before you commit to anything. Or start with the free audit.