// AI RELIABILITY · RAG REFERENCE IMPLEMENTATION

The RAG Decisions, Actually Implemented

A tested RAG pipeline built from the nine decisions in the RAG Architecture Decision Template — not another generic boilerplate with the usual defaults left unexamined. Free to clone and run against your own documents.

// decisions, implemented
01chunkingstructure-aware
02embeddingprovider-agnostic
04retrievalhybrid + RRF
07groundednessnegation-aware
tests27 passing

Same negation-aware groundedness algorithm as Quorel, ported directly — not reinvented. Zero API key required to run the test suite.
// why this isn't another RAG boilerplate

Every choice here traces back to a numbered decision, not a default nobody examined.

Most starter repos make five decisions without looking at any of them, because looking at them is the actual hard part. This one doesn't.

DecisionChoice made here
1. ChunkingStructure-aware — header-first, size-limited fallback
2. EmbeddingProvider-agnostic interface, offline stub by default
3. Vector storeIn-memory by default, pgvector interface included
4. RetrievalHybrid — dense + BM25, fused with reciprocal rank fusion
6. Context assemblyRanked concatenation, explicit seam for your own logic
7. GroundednessNegation-aware citation checking, ported from Quorel
27
tests, all exercising real logic — including the exact negation regression case
0
API keys required to run the test suite offline
6
of 9 template decisions implemented as real, running code
// pricing

Start free. Pay when it's carrying real traffic.

Same model as Quorel, the free tier is the actual pipeline, not a crippled trial. Clone it, run it against your own documents, decide if the decisions fit before you owe anything.

FREE — EVALUATION LICENSE
Clone & Run
$0

Full source, all six implemented decisions, offline stub embedder — everything you need to prove the pattern works before you commit to anything.

  • Structure-aware chunking + hybrid retrieval
  • Negation-aware groundedness checking
  • FastAPI app + 27 passing tests
  • Zero API key required to run tests
  • In-memory vector store
  • Not licensed for production traffic
Clone on GitHub
COMMERCIAL LICENSE
Production
$149 one-time

Everything you need to put this pipeline in front of real documents and real queries, plus guidance for the two production-scale swaps.

  • Everything in Clone & Run
  • License to run in production
  • Walkthrough: wiring a real embedding provider
  • Walkthrough: migrating to pgvector
  • Email support for integration questions
Get the license
DONE WITH YOU
Built for Your Stack
From $1,500

The gaps that are genuinely specific to your environment, closed by hand instead of self-serve.

  • Real embedding provider wired to your model
  • pgvector migration against your actual schema
  • Query transformation, if your eval set shows you need it
  • Retrieval tuning against your own documents
  • Architecture review of your broader RAG system
Start a conversation

clone & run proves the decisions work on your own documents → production gets you shipped fast on your own infra → built for your stack closes what's genuinely yours to solve, not generic to productize.

// the actual question

If your RAG system is already misbehaving and nobody remembers why chunking was set up that way —

it's not whether the decisions were made. It's whether they were made on purpose, and written down anywhere.