Back

GraphRAG Knowledge Engine

Enterprise knowledge graph pipeline

Turning 20K unstructured support tickets into an answerable knowledge base.

Stack

PythonAzure Functions (Durable)Azure Service BusCosmos DBSQL ServerFAISSgpt-4o-minitext-embedding-3-small

Built in a production environment; employer and product names withheld under confidentiality.

What made it hard

The corpus was 20K support tickets written by different people over years, with no clean schema and terminology that varied by author. Keyword search already existed and already failed, so the same problem was being solved from scratch every time because nobody could find the ticket where it had been solved before. There was no ground truth to evaluate retrieval against, so quality had to be established before it could be improved.

Problem

  • Recurring tickets solved from scratch every time
  • Critical knowledge buried across 20K+ documents
  • Keyword search fails on terminology variants
  • No cross-ticket pattern discovery possible
  • Support agents lack contextual answers
  • Hours wasted searching disconnected sources

Solution

  • Two-stage ingest pipeline with graph extraction
  • Graph entities and community detection applied
  • Three parallel vector search strategies fused
  • Pre-generated Q&A cache for common queries
  • FAISS + Cosmos DB hybrid retrieval layer
  • Durable Functions orchestrate bulk processing

Outcome

  • Seconds-to-searchable document indexing achieved
  • Grounded answers with full data citations
  • 90K entities and 150K+ relationships mapped
  • Support knowledge finally reusable at scale
  • 11 Cosmos collections serving live traffic
  • Duplicate ticket resolution dramatically reduced

20K

support documents

90K

entities extracted

150K+

relationships

11

Cosmos collections

Architecture Review

System design · Pipeline · Decisions

System Architecture

How It Works

  1. 1

    Ingestion

    A C# Function App watches the Data Lake. When a new file lands, it drops a Service Bus message — file URL, metadata, RagType. A Python Function App picks it up, pulls the file, and hands it to FileReader, which knows how to read PDF, DOCX, TXT, CSV, JSON, and MD.

  2. 2

    Baseline RAG (seconds)

    The text goes through SmartChunkingUtil, which splits it into overlapping chunks while keeping URLs intact and respecting sentence boundaries. Chunks land in text_units_collection. Another Service Bus message fires the Deep RAG stage.

  3. 3

    Deep RAG (10–60+ minutes)

    gpt-4o-mini reads each chunk and pulls out named entities and typed relationships, with strength scores. It also groups chunks into topics. Every entity gets embedded via text-embedding-3-small. FAISS builds a kNN similarity graph, then hierarchical community detection clusters them recursively. For each community the LLM writes a narrative summary and pre-generates Q&A pairs. Everything persists across 11 Cosmos collections.

  4. 4

    Retrieval (query time)

    The user's question gets embedded, then three Cosmos VectorDistance() searches run in parallel against entities, topics, and questions. Top matches trigger graph expansion. All of that gets formatted into structured tables, injected into the system prompt, and sent to gpt-4o-mini. The model returns a grounded answer with data citations.

Constraints

  • No usable schema: free-text tickets with author-dependent vocabulary for the same concepts
  • No labelled evaluation set existed; retrieval quality had to be measured before it could be tuned
  • Graph enrichment is far too slow to sit in the upload path, but documents had to be searchable immediately
  • LLM cost across 20K documents made a single frontier model per chunk uneconomic
  • Confidential environment: the pipeline had to run entirely inside the client tenancy

Key Decisions

Two-stage pipeline: seconds-to-searchable, minutes-to-graph-enriched

Full graph enrichment takes 10 to 60+ minutes per document. You can't make users wait for that before their doc is even findable. Stage 1 does the cheap work — read, chunk, index — and marks the doc searchable in seconds. Then it fires a Service Bus message that kicks off Stage 2 in the background.

Three parallel vector searches, not one blended index

The obvious first cut is to embed everything into a single collection and query it. The reranker couldn't tell why a chunk matched. Splitting into three collections and fanning out preserves that signal, and merging by similarity afterwards is trivial.

A pre-generated Q&A cache that short-circuits the graph

During Stage 2 the LLM writes out likely question/answer pairs for each chunk. At query time we hit that collection in parallel with the others, and if a question matches strongly we can often skip the whole graph-traversal step.

Two vector systems doing different jobs

Cosmos for online retrieval, FAISS for offline community detection. FAISS runs kNN exactly once to build the entity similarity graph. Cosmos runs the query path where retrieval needs to live right next to the graph data it references.

Service Bus between the C# uploader and the Python pipeline

Different services, different failure modes, different scaling needs. Service Bus in the middle makes the whole thing sane and lets either side redeploy without dropping work.

GraphRAG Knowledge Engine — Architecture

GraphRAG Knowledge Engine architecture

Evaluation

This shipped before a formal evaluation harness was in place — a known gap. Quality was validated by spot-checking answers against known cases and iterating on prompts, chunking, and retrieval weights based on user-reported failures. A proper eval loop is the first thing I'd add if I rebuilt this.

What I'd Change

  • Build the evaluation harness first, not last
  • Entity resolution should have been first, not last
  • Chunking was tuned once and never revisited
  • The Q&A cache stales silently

Get in Touch

Let's build
something great.