The case for GraphRAG
Some questions have no single document that answers them. GraphRAG builds a connected view of the evidence once and reuses it, for broader answers than plain retrieval without rereading everything.
Some questions make a convincing retrieval demo fall apart. What keeps delaying our deals? Which customer problems are becoming a pattern? Where are we making promises the product cannot keep? The answer is spread across calls, tickets, emails and product discussions, and no single document was written to give it. Plain retrieval finds the passages that match the question and misses the ones that explain it. GraphRAG does the organising work once, before anyone asks, and reuses it for every question after. Microsoft’s GraphRAG paper shows that this gives more comprehensive answers than vector retrieval, with far fewer tokens than summarising every document for every question. For a team that keeps asking broad questions about the same business, that is worth building.
A good match can still be a poor answer
In basic retrieval-augmented generation (RAG), the question is used to find relevant text, and the model answers from what comes back. Vector retrieval ranks passages by how similar they are to the question. That works when the question points at particular evidence. It works less well when the question asks for a view across the whole collection, a distinction Microsoft draws in its global search documentation.
Say a sales leader asks why expansion deals keep slipping. Three recent calls mention price. The agent retrieves them and writes an explanation about budget pressure, with citations.
Elsewhere, implementation notes describe an unfinished integration. Support tickets show the same dependency across several accounts. A product discussion says the promised release has moved. None of these passages looks much like the sales leader’s question. Together they point to a different explanation.
The price answer can be faithful to every passage it cites and still miss the pattern. A citation tells you where a sentence came from. It does not tell you the evidence covers the business. Returning more passages helps at the margin, but the system still has no way to know when it has covered enough of the subject to answer a broad question.
Build the connections before you need the answer
GraphRAG adds a preparation step. A language model extracts entities and relationships from the text. The graph is grouped into communities of connected entities, and each community gets a summary, at several levels of detail. Those summaries become reusable context. The indexing documentation describes the stages.
In the expansion example, the entities would be accounts, implementations, integrations and commitments. An account connects to an implementation, the implementation depends on an integration and several other accounts depend on that integration too. Those links let the system group the evidence around the shared problem, even when the sources describe it in different words.
For a broad question, global search works through batches of community summaries, turns each batch into an intermediate answer, then filters and combines the useful points into a final one. The question is answered from an organised view of the whole collection, not from whatever a search happened to return.
What I find compelling is the split. Organising the evidence is done once and reused. Answering is done per question, against that preparation. The next question starts from the same structure instead of asking the model to rebuild the picture.
The graph matters because it decides what gets summarised together. Putting documents in a graph database would not do this. The work is in extracting the relationships, grouping the evidence and writing the summaries. The bet is that relationships are a good way to group company evidence, because a shared product dependency often matters more than which document a fact appears in. That is the part to test against simpler ways of preparing summaries.
The argument is coverage first, compression second
Two comparisons matter. Against vector retrieval, the question is whether a wider view helps. Against summarising all the source text, the question is what the graph adds once both systems see the whole collection.
The chart shows root-level GraphRAG (C0), the same configuration as the token comparison below. Its comprehensiveness score was about 72% against vector retrieval on both datasets. Against full-text summarisation, its advantage was not statistically significant (Table 6).
For a question about recurring customer problems, coverage is what to check: does the answer include what the team needs to see? A short answer is no use if it leaves out the dependency affecting half the accounts under review.
The preparation also cuts how much text each question needs. C0 used over 97% fewer context tokens than summarising all the source text directly.
The saving matters most when the same collection answers many questions. Preparation has a cost, so judge it over the life of the workload: building the index, keeping it current and answering the questions people actually ask. A one-off summary and a weekly review deserve different sums.
The study has limits. It used podcast and news corpora, not company workflows. Its claim-based follow-up found no significant difference in comprehensiveness or diversity between GraphRAG and full-text summarisation. Vector retrieval gave more direct answers, and the paper does not show that GraphRAG reduces hallucination (results and limitations).
So the study makes a strong case for broader coverage and a practical case for getting it through reusable summaries. It is a reason to test GraphRAG on company evidence, not a result that settles the test.
Use GraphRAG where the question crosses the documents
Start with a question the team keeps answering by opening several systems and rebuilding the same relationships: recurring implementation blockers, themes across lost deals, shared dependencies behind customer requests. When the hard part is assembling the picture, a graph has something to improve.
Take past examples and ask the people who know the work to judge the answers from plain retrieval, whole-collection summarisation and GraphRAG. Which important themes did each one miss? Are the connections supported by the sources? Can the reader get back to the evidence? Is the improvement worth the cost of keeping the index current?
For a lookup, such as a clause in a known contract, use plain retrieval. For an exact count, such as overdue invoices, query the records. And before an agent acts on any of it, it still needs current state and permissions, which is where a world model for knowledge work comes in.
Retrieval answers the questions that point at a document. GraphRAG is for the ones whose answer sits between them.