ð ðð¹ðð¶-ðµðŒðœ ð¥ðð ðð¶ððµðŒðð ð® ðŽð¿ð®ðœðµ ð±ð®ðð®ð¯ð®ðð² Multi-hop retrieval is often ...

TL;DR · AI æèŠ
Milvuså¢éåŒåäºäžç§æ éåŸæ°æ®åºçå€è·³RAGæ¹æ³ïŒéè¿åéååšåID亀ååŒçšå®ç°é«ææšçïŒéäœAPIææ¬å¹¶æåæ§èœã
æ žå¿èŠç¹
- Vector Graph RAG䜿çšäžäžªMilvuséåååšå®äœãå ³ç³»åæ®µèœïŒéè¿ID亀ååŒçšæå»ºåŸç»æã
- è¯¥æ¹æ³æ¯æ¥è¯¢ä» é2次LLMè°çšïŒæ¯IRCoTåAgentic RAGåå°50%以äžçAPIææ¬ã
- åšMuSiQueçåºåæµè¯äžèŸŸå°85.8%çRecall@5ïŒäŒäºHippoRAG 2ç85.1%ã
ç»ææçº²
æç« èå¿«é跳蜬ã
- §åŒèš
ä»ç»å€è·³RAGéåžžäŸèµåŸæ°æ®åºïŒäœMilvuså¢éå°è¯ä» çšMilvuså®ç°ã
Vector Graph RAGéè¿äžäžªéåååšæ°æ®ïŒå©çšID亀ååŒçšæå»ºåŸç»æã
- âºå®ç°ç»è
å®äœãå ³ç³»ãæ®µèœçååšæ¹åŒåIDåŒçšå®ç°åŸç»æã
åå°LLMè°ç𿬡æ°ïŒéäœAPIææ¬ïŒæåååºé床ã
- âºå®éªç»æ
åšå€äžªåºåæµè¯äžè¡šç°äŒåŒïŒRecall@5蟟85.8%ã
éè¿pipå®è£ ïŒæ¯ææ¬å°Milvus Liteã
æç»Žå¯ŒåŸ
çšäžåŒ åŸçæž äž»é¢ä¹éŽçå ³ç³»ã
æ¥çå€§çº²ææ¬ïŒæ éç¢ / æ JS å奜ïŒ
- Vector Graph RAG
- æ žå¿æºå¶
- ID亀ååŒçšæå»ºåŸç»æ
- æ§èœäŒå¿
- 2次LLMè°çš/æ¥è¯¢
- 60% APIææ¬éäœ
- å®éªç»æ
- 85.8% Recall@5
éå¥ / Highlights
åŒåŸæ¶èäžå享çå ³é®å¥ã
Vector Graph RAG stores everything in three Milvus collections with ID cross-references.
â 第 2 段
That pipeline makes 2 LLM calls per query... 60% lower API cost and 2-3x faster responses.
â 第 3 段
85.8% average Recall@5... against 85.1% for HippoRAG 2.
â 第 4 段
Milvus on X: "ð ðð¹ðð¶-ðµðŒðœ ð¥ðð ðð¶ððµðŒðð ð® ðŽð¿ð®ðœðµ ð±ð®ðð®ð¯ð®ðð² Multi-hop retrieval is often solved by adding a separate graph store. We wanted to see how far we could get with Milvus alone â so we built an open-source library and found out. Vector Graph RAG achieves multi-hop reasoning using only Milvus. Neo4j, Cypher queries, and the second system once needed to operate are all things of the past (in this scenario). Vector Graph RAG is built upon the fact that knowledge graph relations are just text. An example, (metformin, is the first-line drug for, type 2 diabetes) is a directed edge in a graph database â but it's also a sentence you can embed and store in Milvus, alongside entities and source passages. ð©ð²ð°ððŒð¿ ðð¿ð®ðœðµ ð¥ðð ðððŒð¿ð²ð ð²ðð²ð¿ðððµð¶ð»ðŽ ð¶ð» ððµð¿ð²ð² ð ð¶ð¹ððð ð°ðŒð¹ð¹ð²ð°ðð¶ðŒð»ð ðð¶ððµ ðð ð°ð¿ðŒðð-ð¿ð²ð³ð²ð¿ð²ð»ð°ð²ð:⢠ðð»ðð¶ðð¶ð²ð â deduplicated nodes, each carrying a list of the relation IDs they participate in ⢠ð¥ð²ð¹ð®ðð¶ðŒð»ð â embedded triples pointing to the entity IDs and source passage IDs on each side ⢠ð£ð®ððð®ðŽð²ð â original document chunks with back-references to the extracted entities and relations Those ID references are the graph structure. Subgraph expansion follows them to surface bridge entities that the question never mentions â the step that makes multi-hop reasoning work. Then a single LLM reranking pass filters the expanded candidate pool down to what actually answers the question, and one generation call produces the answer from the full source passages. That pipeline makes ð® ððð ð°ð®ð¹ð¹ð ðœð²ð¿ ðŸðð²ð¿ð (rerank + generate), compared to 3-5 for IRCoT and 5-10+ for Agentic RAG. Against a 5-call iterative baseline, that works out to roughly ð²ð¬% ð¹ðŒðð²ð¿ ðð£ð ð°ðŒðð ð®ð»ð± ð®-ð¯ð ð³ð®ððð²ð¿ ð¿ð²ððœðŒð»ðð²ð, with predictable latency instead of spikes when an agent decides to loop again. On the benchmarks: ðŽð³.ðŽ% ð®ðð²ð¿ð®ðŽð² ð¥ð²ð°ð®ð¹ð¹@ð± across MuSiQue, HotpotQA, and 2WikiMultiHopQA â against ðŽð³.ð% for HippoRAG 2, under the same evaluation setup, with no graph database and no ColBERTv2. ððŒð ð±ðŒ ððŒð ð¶ð»ððð®ð¹ð¹ ððµð¶ð? You only need to type"pip install vector-graph-rag". It defaults to Milvus Lite â a local .db file, so you don't need to configure anything. ððŒð¿ ðºðŒð¿ð² ð¶ð»ð³ðŒð¿ðºð®ðð¶ðŒð», ðð²ð² ð¶ðð ðŽð¶ððµðð¯ ðœð®ðŽð²: https://t.co/uyBstuGBcq If your corpus is knowledge-dense â legal, biomedical, financial â and your questions routinely cross 2-4 document boundaries, this is the architecture worth testing first." / X
Milvus
@milvusio
ð ðð¹ðð¶-ðµðŒðœ ð¥ðð ðð¶ððµðŒðð ð® ðŽð¿ð®ðœðµ ð±ð®ðð®ð¯ð®ðð² Multi-hop retrieval is often solved by adding a separate graph store. We wanted to see how far we could get with Milvus alone â so we built an open-source library and found out. Vector Graph RAG achieves multi-hop reasoning using only Milvus. Neo4j, Cypher queries, and the second system once needed to operate are all things of the past (in this scenario). Vector Graph RAG is built upon the fact that knowledge graph relations are just text. An example, (metformin, is the first-line drug for, type 2 diabetes) is a directed edge in a graph database â but it's also a sentence you can embed and store in Milvus, alongside entities and source passages. ð©ð²ð°ððŒð¿ ðð¿ð®ðœðµ ð¥ðð ðððŒð¿ð²ð ð²ðð²ð¿ðððµð¶ð»ðŽ ð¶ð» ððµð¿ð²ð² ð ð¶ð¹ððð ð°ðŒð¹ð¹ð²ð°ðð¶ðŒð»ð ðð¶ððµ ðð ð°ð¿ðŒðð-ð¿ð²ð³ð²ð¿ð²ð»ð°ð²ð:⢠ðð»ðð¶ðð¶ð²ð â deduplicated nodes, each carrying a list of the relation IDs they participate in ⢠ð¥ð²ð¹ð®ðð¶ðŒð»ð â embedded triples pointing to the entity IDs and source passage IDs on each side ⢠ð£ð®ððð®ðŽð²ð â original document chunks with back-references to the extracted entities and relations Those ID references are the graph structure. Subgraph expansion follows them to surface bridge entities that the question never mentions â the step that makes multi-hop reasoning work. Then a single LLM reranking pass filters the expanded candidate pool down to what actually answers the question, and one generation call produces the answer from the full source passages. That pipeline makes ð® ððð ð°ð®ð¹ð¹ð ðœð²ð¿ ðŸðð²ð¿ð (rerank + generate), compared to 3-5 for IRCoT and 5-10+ for Agentic RAG. Against a 5-call iterative baseline, that works out to roughly ð²ð¬% ð¹ðŒðð²ð¿ ðð£ð ð°ðŒðð ð®ð»ð± ð®-ð¯ð ð³ð®ððð²ð¿ ð¿ð²ððœðŒð»ðð²ð, with predictable latency instead of spikes when an agent decides to loop again. On the benchmarks: ðŽð³.ðŽ% ð®ðð²ð¿ð®ðŽð² ð¥ð²ð°ð®ð¹ð¹@ð± across MuSiQue, HotpotQA, and 2WikiMultiHopQA â against ðŽð³.ð% for HippoRAG 2, under the same evaluation setup, with no graph database and no ColBERTv2. ððŒð ð±ðŒ ððŒð ð¶ð»ððð®ð¹ð¹ ððµð¶ð? You only need to type"pip install vector-graph-rag". It defaults to Milvus Lite â a local .db file, so you don't need to configure anything. ððŒð¿ ðºðŒð¿ð² ð¶ð»ð³ðŒð¿ðºð®ðð¶ðŒð», ðð²ð² ð¶ðð ðŽð¶ððµðð¯ ðœð®ðŽð²:
github.com/zilliztech/vecâŠ
If your corpus is knowledge-dense â legal, biomedical, financial â and your questions routinely cross 2-4 document boundaries, this is the architecture worth testing first.
3:00 PM · Jul 23, 2026
174
Views
1
3