Case study
Citation Desk
Answers grounded in your documents, with page-level citations.
Reply with the source passage and a retrieval route
- 0.4s
- Cited lookup
- 3
- Retrieval routes
- 100%
- Answers with sources
The problem with uncited answers
Teams were pasting documents into a chat window and treating the reply as a summary. When a sentence was wrong, nobody could tell whether retrieval missed the page or the model invented a detail.
Citation Desk is the placeholder case study for a source-grounded assistant. Replace this narrative with the system you actually shipped.
Engineering approach
Documents are split into overlapping windows, embedded, and stored with page and line metadata. At question time a router chooses document search, a web fallback, or a refusal.
The writer receives only the selected spans. If the top score is below a threshold, the product says it cannot support an answer.
- Chunking preserves page boundaries so citations stay precise
- A router keeps web results out of private-document questions
- The prompt forbids claims that are not present in the spans
Architecture
The interface streams tokens and a side panel of pipeline events. The API owns indexing, retrieval, and generation so the client never holds document vectors.
Results
In sample evaluations, cited answers were reviewable in a single click, and unsupported questions failed closed. Swap these results for your own measurements.
Technical implementation
Next.js chat UI, a FastAPI streaming API, a chunk-and-embed indexer, and a pgvector store. The generator is only allowed to answer from retrieved spans.
- Next.js
- FastAPI
- LangGraph
- pgvector
- Python
Key features
- Passage-level citations
- Route trace for every answer
- Upload and re-index without a redeploy
- Refusal when retrieval confidence is low
Written by
Anuoluwa Olutayo
AI Engineer