Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Read by thought-leaders and decision-makers around the world. Phone Number: +1-650-246-9381 Email: pub@towardsai.net
228 Park Avenue South New York, NY 10003 United States
Website: Publisher: https://towardsai.net/#publisher Diversity Policy: https://towardsai.net/about Ethics Policy: https://towardsai.net/about Masthead: https://towardsai.net/about
Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Founders: Roberto Iriondo, , Job Title: Co-founder and Advisor Works for: Towards AI, Inc. Follow Roberto: X, LinkedIn, GitHub, Google Scholar, Towards AI Profile, Medium, ML@CMU, FreeCodeCamp, Crunchbase, Bloomberg, Roberto Iriondo, Generative AI Lab, Generative AI Lab VeloxTrend Ultrarix Capital Partners Denis Piffaretti, Job Title: Co-founder Works for: Towards AI, Inc. Louie Peters, Job Title: Co-founder Works for: Towards AI, Inc. Louis-François Bouchard, Job Title: Co-founder Works for: Towards AI, Inc. Cover:
Towards AI Cover
Logo:
Towards AI Logo
Areas Served: Worldwide Alternate Name: Towards AI, Inc. Alternate Name: Towards AI Co. Alternate Name: towards ai Alternate Name: towardsai Alternate Name: towards.ai Alternate Name: tai Alternate Name: toward ai Alternate Name: toward.ai Alternate Name: Towards AI, Inc. Alternate Name: towardsai.net Alternate Name: pub.towardsai.net
5 stars – based on 497 reviews

Frequently Used, Contextual References

TODO: Remember to copy unique IDs whenever it needs used. i.e., URL: 304b2e42315e

Resources

Free: 6-day Agentic AI Engineering Email Guide.
Learnings from Towards AI's hands-on work with real clients.
Vector Search Is Not Enough: The Blueprint of GraphRAG
Latest   Machine Learning

Vector Search Is Not Enough: The Blueprint of GraphRAG

Last Updated on October 6, 2026 by Editorial Team

Author(s): Asma Houimli

Originally published on Towards AI.

Vector Search Is Not Enough: The Blueprint of GraphRAG

Vector Search Is Not Enough: The Blueprint of GraphRAG

Retrieval-Augmented Generation (RAG) has emerged as an effective method for improving the factual accuracy and grounding of Large Language Models (LLMs) by integrating external knowledge during the inference process.

However, recent research has shifted its focus to graph-structured grounding data. This change is made to address several limitations in RAG. For instance, RAG cannot extract relevant information unless it is explicitly mentioned within the text chunks. If certain relationships among the entities present in the text are not specifically stated, RAG is unable to capture these relationships effectively.

For example, if we consider the question:

“What are the main findings of the study conducted by Dr. Smith in 2020 on lung cancer treatment? ”

In a standard RAG system, the sentence mentioning “Dr. Smith” might be in one chunk, the publication year in another, the topic and main finding in yet another. Because embeddings treat chunks in isolation, the model may retrieve a few fragments, miss key connections, and therefore fail to generate a coherent or accurate answer.

To address this, GraphRAG was introduced in 2024; this shift is intended to capture the rich relational context that often implicitly exists in complex fields. As a result, the term “GraphRAG” has emerged to describe the combination of knowledge graphs and RAG.

Microsoft first introduced the concept in their paper titled "From Local to Global: A GraphRAG Approach to Query-Focused Summarization." In their paper, the authors present a GraphRAG method that involves creating a graph of community summaries at various levels of granularity. These summaries are then used as relevant context to provide input to the LLM, ultimately helping to generate the final answer.

GraphRAG

Similar to the basic RAG pipeline, GraphRAG systems consist essentially of two components:

  • A Retriever: As indicated by its name, it extracts data from the knowledge graph that may be relevant or meaningful for answering questions.
  • A Generator: uses the data extracted by the retriever to formulate the final answer to the query.

There are various implementations of these components, as well as techniques that introduce other components.

The generator usually reformulates the retrieved context into a natural language answer, so it has minimal control over which information is used. As a result, the quality of the final answer largely depends on how effective the retriever is. If a retriever is poorly designed, it may not surface the most relevant information or might provide noisy or redundant context, which can mislead the generator or diminish the quality of its response.

GraphRAG

It is important to acknowledge that, in addition to the retriever and generator, constructing the knowledge graph is a vital step in the process. There are various strategies for creating the source graph. Some systems extract entities and relationships through rule-based or NLP-driven information extraction pipelines, while others leverage the capabilities of LLMs to build the graph.

In this article, we will dive deeper into the world of GraphRAG: From graph creation to query answering.

What is a knowledge graph?

A knowledge graph, also sometimes called a ‘semantic network’, is a structured representation of data in which entities are represented as nodes and the relationships between them as edges. This format is particularly useful for integrating and reasoning about diverse and often unstructured data, such as texts, web pages, and research articles.

  • Nodes (also called vertices): represent real-world or abstract entities, such as a person, place, organization, event, or concept. Each node can carry additional properties (attributes) that describe it in more detail. For example, a node representing a Person might include attributes like name, age, and occupation.
  • Edges: represent the connection between two nodes, illustrating semantic relationships between nodes. For example works_at, born_in, or part_of.

In the context of GraphRAG, we primarily discuss text-based heterogeneous labeled graphs. This article focuses on these types of graphs and does not include homogeneous, unlabeled graphs.

Example of a knowledge graph (citation network)

Ontology

When discussing knowledge graphs, the term "ontology" is likely to come up. There is still some debate regarding whether an ontology is different from a knowledge graph. An ontology typically serves as the blueprint or rulebook for a knowledge graph. It includes a formal model that defines entities, relationships, attributes, and logical rules for a specific domain. Essentially, it acts as a schema or grammar, explaining what concepts mean and how they can relate to one another.

Graphs are typically stored and queried using graph databases like Neo4j, TigerGraph, Amazon Neptune, etc., which support expressive query languages like Cypher and SPARQL.

Knowledge Graph Construction Methods

  • Manual construction

This approach is the most challenging and time-consuming, but it offers the highest quality, as the graph is constructed through human annotation. For instance, creating a biological knowledge graph would require the expertise of biologists who manually curate entities and relationships based on trusted sources like UMLS (Unified Medical Language System), DrugBank, or NPASS.

  • Rule-based construction

To reduce the need for manual annotation, many traditional methods rely on rule-based parsers to extract relevant nodes and relationships from raw text. For example, Named Entity Recognition (NER) tools, such as the traditional SpaCy or the more modern GliNER, can identify entities like people, organizations, and locations within a sentence. Similarly, customized rules or pattern-based extractors can be employed to extract relationships.

  • LLM-based construction

Recent research has focused on using LLMs to construct graphs from a corpus of unstructured documents, leveraging their ability to extract entities and relationships from raw text. This method requires minimal to no human intervention. However, it can be an expensive approach and may produce unpredictable results since we cannot fully control the behavior of the LLM.

For example, in the Microsoft paper titled "From Local to Global: A Graph RAG Approach to Query-Focused Summarization," they discussed how LLMs can be utilized to create community graphs. Several popular solutions for automatic knowledge graph building are available today, such as Knowledge Graph Index from LlamaIndex and the Neo4j LLM knowledge graph builder, which allows users to have a bit of control over the structure and the content of the generated graph by defining the schema of the graph that will be created (the entities and relationships allowed).

Other tools specifically designed for codebases include Graphify and CodeGraph. These tools are used to build queryable knowledge graphs from code, documentation, papers, and diagrams, assisting coding agents in navigating and understanding large codebases more effectively.

GraphRAG pipelines

As mentioned above, the effectiveness of the GraphRAG pipeline lies in the retriever component, which has recently been integrated with LLMs to enhance their ability to access and incorporate external knowledge, improving response accuracy. There are various techniques and implementations of retrievers, and we will review some of them.

A. Text-attributed knowledge graph / Domain graph

Become a Medium member

The required knowledge graph for the following strategies is text-attributed/ domain graph, where we represent real-world entities and the relationships between them, as shown below:

Example of a domain graph
  1. Graph Traversal: Cypher-based

After creating the knowledge graph and storing it in a graph-based database like Neo4j, we need to translate user questions into queries that can be executed on the database to retrieve structured data.

Graph traversal: Cypher-based retrieval
  • Cypher Template: A basic approach involves using predefined queries created by domain experts, which can be mapped to user questions. When a user poses a question, an LLM determines which Cypher template to use. The LLM may also extract parameters from the user’s question and insert them into the chosen template. This query is then executed on the database, and the results are sent back to the LLM to generate a response.

However, this method is only effective when the types of questions users are likely to ask about the Domain Graph are already known, and corresponding templates have been prepared in advance.

  • Dynamic Cypher Generation: This method seeks to improve upon the previous approach. Many user questions often include multiple filters, though they may not always be the same.

For example: we may encounter several related questions such as

  • Which research papers has Ian Goodfellow written?
  • Which research papers has Ian Goodfellow written between 2019 and 2023?
  • Which research papers has Ian Goodfellow written between 2019 and 2023 that …?

The possibilities are endless. We wouldn’t want to create a Cypher template for each of these questions. The solution is to dynamically generate Cypher queries based on the parameters actually given in the user question.

Given a user question, an LLM decides which of the Cypher templates to use. Multiple templates can also be used in a chain or loop, which leads to an agentic query system.

However, even with this improvement, the range of questions remains limited by the available Cypher templates.

  • Text2Cypher: The two previously discussed methods are both limited by the queries/ query snippets that are defined during implementation.

This method was presented by Neo4j in their paper “Text2Cypher: Bridging Natural Language and Graph Databases”, in which the user question and the graph schema or ontology are combined in a prompt and provided to the LLM to translate it into a Cypher query, which is then executed against the database. Once the Cypher query is executed, the retrieved subgraph is used to generate the final answer.

This pattern is highly flexible, as there are no predefined queries, meaning that in theory, the language model can generate any query that fits the provided schema. However, this approach is not completely reliable due to the limitations in the foundational models’ understanding of the Cypher query language. The situation can be improved by fine-tuning the models for better query generation, as discussed in the paper.

2. Agent-based retrieval

Other researchers in the field went for more general approaches by using agents. For instance, in their paper “Graph Chain-of-Thought: Augmenting Large Language Models by Reasoning on Graphs”, Jin et al. presented Graph Chain-of-Thought (Graph-COT), an iterative framework that defines tools, or simply functions, that retrieve nodes, attributes, or traverse the graph by visiting neighboring nodes.

The process involves providing the LLM with both a query and a textual description of the schema or ontology of the knowledge graph. The LLM then reasons iteratively on the graph and utilizes the appropriate tools to arrive at an answer, or it stops once a specific limit on the number of iterations is reached. This approach encourages the decomposition of complex tasks into several executable steps.

Graph Chain of Thought

This framework succeeded at tackling local-focused queries, since it processes the graph by concentrating on local interactions rather than considering the whole graph at once.

However, it can be costly and slower than the methods described previously, especially with complex questions, since we need to have multiple small interactions that necessitate reasoning by the LLM.

B. Lexical graph

A lexical graph maps the structural organization, documents, chunks, and entities rather than just considering real-world entities:

Lexical graph with extracted entities
  • Graph-enhanced vector search

This representation connects real-world entities across data chunks, allowing us to retrieve these relationships alongside a vector search. This approach provides additional context about the entities referred to in the chunks.

When a user submits a query, it is embedded, and the top k most relevant chunks (configured in advance) are retrieved. However, unlike basic RAG approaches, we go further by conducting a traversal that starts with the found chunks. This traversal aims to gather more relevant context by showcasing the entities found there and their relationships.

This method is beneficial for obtaining a richer context than what a standard vector search would provide. The additional traversal uncovers interactions among the entities within the data, revealing much more comprehensive information. This is the type of graph that Neo4j’s knowledge graph builder creates.

C. Community summaries graph

This is the Microsoft GraphRAG, the graph in general will be something like this:

Community summaries graph

Certain questions can be asked about the entire dataset, rather than just focusing on specific chunks or parts of the knowledge graph. These questions seek global insights. The patterns mentioned earlier are not suitable for addressing such "global" questions.

To create this type of community graph, we need to extract entities and relationships while forming hierarchical communities within the domain graph. For each community, an LLM summarizes the entity and relationship information into community summaries.

Given a question, we can adopt a simple approach for finding answers. First, we can look for higher-level community summaries to see if they provide the information we need. If those summaries are insufficient, we can then delve into lower-level community summaries and so on until we find the answer.

Alternatively, we can use a strategy called DRIFT search (Dynamic Reasoning and Inference with Flexible Traversal). This method is based on Microsoft’s GraphRAG technique and utilizes a multi-stage approach that connects high-level global context with detailed local information. It begins with a broad vector-based community search and then formulates additional questions based on the initial results to conduct more localized searches. After collecting all the results, we re-rank them and utilize the best ones to generate the final answer.

This pattern is useful for addressing global questions; however, setting up the required graph involves significant effort due to the many steps involved: entity and relationship extraction, community detection, and community summarization. Relying entirely on an LLM to perform these tasks can be costly and may compromise the quality of the resulting graph.

In conclusion, recent research has increasingly focused on using knowledge graphs as foundational data sources. Today, we see many products and applications that leverage knowledge graphs to improve context management and enhance the memory capabilities of agents. However, as previously mentioned, creating these knowledge graphs from documents can be expensive. Additionally, selecting the wrong retrieval approach can lead to inefficiencies that are both slow and costly. Therefore, the choice of method largely depends on the specific use case, as well as the desired latency and overall cost.

Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.

Published via Towards AI


Towards AI Academy

We Build Enterprise-Grade AI. We'll Teach You to Master It Too.

15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.

Start free — no commitment:

→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day

→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages

Our courses:

→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.

→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.

→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.

Note: Article content contains the views of the contributing authors and not Towards AI.