# GraphRAG Knowledge Engine

> Enterprise knowledge graph pipeline

Turning 20K unstructured support tickets into an answerable knowledge base.

- **Page:** https://siddharthdeshpande.com/projects/graphrag
- **Built at:** Confidential
- **Note:** Built in a production environment; employer and product names withheld under confidentiality.
- **Tech stack:** Python, Azure Functions (Durable), Azure Service Bus, Cosmos DB, SQL Server, FAISS, gpt-4o-mini, text-embedding-3-small

## Impact

- **20K** — support documents
- **90K** — entities extracted
- **150K+** — relationships
- **11** — Cosmos collections

## Problem

- Recurring tickets solved from scratch every time
- Critical knowledge buried across 20K+ documents
- Keyword search fails on terminology variants
- No cross-ticket pattern discovery possible
- Support agents lack contextual answers
- Hours wasted searching disconnected sources

## Solution

- Two-stage ingest pipeline with graph extraction
- Graph entities and community detection applied
- Three parallel vector search strategies fused
- Pre-generated Q&A cache for common queries
- FAISS + Cosmos DB hybrid retrieval layer
- Durable Functions orchestrate bulk processing

## Outcome

- Seconds-to-searchable document indexing achieved
- Grounded answers with full data citations
- 90K entities and 150K+ relationships mapped
- Support knowledge finally reusable at scale
- 11 Cosmos collections serving live traffic
- Duplicate ticket resolution dramatically reduced

## How it works

1. **Ingestion** — A C# Function App watches the Data Lake. When a new file lands, it drops a Service Bus message — file URL, metadata, RagType. A Python Function App picks it up, pulls the file, and hands it to FileReader, which knows how to read PDF, DOCX, TXT, CSV, JSON, and MD.
2. **Baseline RAG (seconds)** — The text goes through SmartChunkingUtil, which splits it into overlapping chunks while keeping URLs intact and respecting sentence boundaries. Chunks land in text_units_collection. Another Service Bus message fires the Deep RAG stage.
3. **Deep RAG (10–60+ minutes)** — gpt-4o-mini reads each chunk and pulls out named entities and typed relationships, with strength scores. It also groups chunks into topics. Every entity gets embedded via text-embedding-3-small. FAISS builds a kNN similarity graph, then hierarchical community detection clusters them recursively. For each community the LLM writes a narrative summary and pre-generates Q&A pairs. Everything persists across 11 Cosmos collections.
4. **Retrieval (query time)** — The user's question gets embedded, then three Cosmos VectorDistance() searches run in parallel against entities, topics, and questions. Top matches trigger graph expansion. All of that gets formatted into structured tables, injected into the system prompt, and sent to gpt-4o-mini. The model returns a grounded answer with data citations.

## Key decisions

- **Two-stage pipeline: seconds-to-searchable, minutes-to-graph-enriched** — Full graph enrichment takes 10 to 60+ minutes per document. You can't make users wait for that before their doc is even findable. Stage 1 does the cheap work — read, chunk, index — and marks the doc searchable in seconds. Then it fires a Service Bus message that kicks off Stage 2 in the background.
- **Three parallel vector searches, not one blended index** — The obvious first cut is to embed everything into a single collection and query it. The reranker couldn't tell why a chunk matched. Splitting into three collections and fanning out preserves that signal, and merging by similarity afterwards is trivial.
- **A pre-generated Q&A cache that short-circuits the graph** — During Stage 2 the LLM writes out likely question/answer pairs for each chunk. At query time we hit that collection in parallel with the others, and if a question matches strongly we can often skip the whole graph-traversal step.
- **Two vector systems doing different jobs** — Cosmos for online retrieval, FAISS for offline community detection. FAISS runs kNN exactly once to build the entity similarity graph. Cosmos runs the query path where retrieval needs to live right next to the graph data it references.
- **Service Bus between the C# uploader and the Python pipeline** — Different services, different failure modes, different scaling needs. Service Bus in the middle makes the whole thing sane and lets either side redeploy without dropping work.

## Evaluation

This shipped before a formal evaluation harness was in place — a known gap. Quality was validated by spot-checking answers against known cases and iterating on prompts, chunking, and retrieval weights based on user-reported failures. A proper eval loop is the first thing I'd add if I rebuilt this.

## What I'd change

- **Build the evaluation harness first, not last** — Without one, every tuning decision becomes a guess dressed up as intuition. Even a small held-out set of 30–50 real queries with reference answers would have caught issues months earlier.
- **Entity resolution should have been first, not last** — "Access Panel X200" and "AP-X200" all need to collapse to the same node, otherwise the graph fragments and community detection produces nonsense.
- **Chunking was tuned once and never revisited** — A dense PDF manual and a 3-line CSV row almost certainly want different chunking strategies; instead they got the same one.
- **The Q&A cache stales silently** — New products ship, new incident patterns emerge, and the cache drifts out of relevance with nothing to flag it. A basic hit-rate metric would surface this immediately.
