Knowledge Graphs Enhance Retrieval Accuracy in RAG Systems
Researchers from the Vector Institute have demonstrated that integrating knowledge graphs into retrieval-augmented generation systems significantly improves accuracy and multi-hop reasoning for complex data. The study introduces an Entity-Based KG-RAG architecture designed to overcome the structural limitations of traditional vector-based RAG, which treats documents as isolated text chunks and frequently misses critical relational context. Standard RAG relies on semantic similarity matching between query embeddings and indexed document fragments. While efficient for straightforward queries, this method struggles with tasks requiring cross-document synthesis, such as connecting financial metrics across different reporting periods or preserving hierarchical document metadata. The researchers developed a pipeline that constructs a structured knowledge graph from unstructured text using large language models to extract entities and relationships. A key innovation involves preserving document-level context by hyphenating extracted entities with their source titles before embedding, preventing hierarchical data loss during processing. The retrieval workflow operates in two phases. First, the system identifies top-matching entities based on query similarity. Second, it explores the local subgraph surrounding these entities to uncover interconnected facts, followed by a chunk-voting mechanism that selects relevant passages based on entity frequency and semantic alignment. This hybrid approach combines structural graph traversal with vector matching, enabling navigation of complex relational pathways. Evaluations conducted on a specialized dataset of SEC 10-Q quarterly reports from major technology firms confirmed the method’s effectiveness. Against synthetic multi-hop financial queries, the Entity-Based KG-RAG system increased answer accuracy from 40 to 55 percent using GPT-4o, and from 36.36 to 56 percent using GPT-4o-mini. The smaller model showed disproportionate gains, indicating that structural graph data effectively compensates for the reasoning limitations of less capable language models. Error analysis revealed that the graph-enhanced system successfully recovered numerous queries missed by baseline approaches, while maintaining high precision on questions the baseline answered correctly. Latency measurements confirmed that graph traversal added negligible overhead to standard retrieval workflows. Alternative graph-based methods, including Cypher-driven retrieval and community-detection frameworks like Microsoft’s GraphRAG, were also evaluated. While promising for specific use cases, these alternatives demonstrated lower accuracy on the financial dataset due to brittle query generation and suboptimal clustering. The Entity-Based approach remained the most robust, peaking when processing between 30 and 40 top-matching entities per query. The findings position knowledge graphs as a critical advancement for RAG systems deployed in domains requiring precise relational reasoning, including finance, legal analysis, and technical documentation. By explicitly modeling entity connections, KG-RAG architectures deliver more transparent and accurate responses without compromising processing speed. The complete implementation and evaluation scripts have been released as open-source code, providing a foundation for structurally aware AI retrieval development.
