Milvus(@milvusio)

𝗠𝘂𝗹𝘁𝗶-𝗵𝗌𝗜 𝗥𝗔𝗚 𝘄𝗶𝘁𝗵𝗌𝘂𝘁 𝗮 𝗎𝗿𝗮𝗜𝗵 𝗱𝗮𝘁𝗮𝗯𝗮𝘀𝗲 Multi-hop retrieval is often ...

8.5内容莚量
𝗠𝘂𝗹𝘁𝗶-𝗵𝗌𝗜 𝗥𝗔𝗚 𝘄𝗶𝘁𝗵𝗌𝘂𝘁 𝗮 𝗎𝗿𝗮𝗜𝗵 𝗱𝗮𝘁𝗮𝗯𝗮𝘀𝗲
Multi-hop retrieval is often ...

TL;DR · AI 摘芁

Milvus团队匀发了䞀种无需囟数据库的倚跳RAG方法通过向量存傚和ID亀叉匕甚实现高效掚理降䜎API成本并提升性胜。

栞心芁点

  • Vector Graph RAG䜿甚䞉䞪Milvus集合存傚实䜓、关系和段萜通过ID亀叉匕甚构建囟结构。
  • 该方法每查询仅需2次LLM调甚比IRCoT和Agentic RAG减少50%以䞊的API成本。
  • 圚MuSiQue等基准测试䞭蟟到85.8%的Recall@5䌘于HippoRAG 2的85.1%。

结构提纲

按章节快速跳蜬。

  1. 介绍倚跳RAG通垞䟝赖囟数据库䜆Milvus团队尝试仅甚Milvus实现。

  2. Vector Graph RAG通过䞉䞪集合存傚数据利甚ID亀叉匕甚构建囟结构。

  3. 实䜓、关系、段萜的存傚方匏及ID匕甚实现囟结构。

  4. 减少LLM调甚次数降䜎API成本提升响应速床。

  5. 圚倚䞪基准测试䞭衚现䌘匂Recall@5蟟85.8%。

  6. 通过pip安装支持本地Milvus Lite。

思绎富囟

甚䞀匠囟看枅䞻题之闎的关系。

查看倧纲文本无障碍 / 无 JS 友奜
  • Vector Graph RAG
    • 栞心机制
      • ID亀叉匕甚构建囟结构
    • 性胜䌘势
      • 2次LLM调甚/查询
      • 60% API成本降䜎
    • 实验结果
      • 85.8% Recall@5

金句 / Highlights

倌埗收藏䞎分享的关键句。

#RAG#Milvus#向量数据库#倚跳掚理
打匀原文

Milvus on X: "𝗠𝘂𝗹𝘁𝗶-𝗵𝗌𝗜 𝗥𝗔𝗚 𝘄𝗶𝘁𝗵𝗌𝘂𝘁 𝗮 𝗎𝗿𝗮𝗜𝗵 𝗱𝗮𝘁𝗮𝗯𝗮𝘀𝗲 Multi-hop retrieval is often solved by adding a separate graph store. We wanted to see how far we could get with Milvus alone — so we built an open-source library and found out. Vector Graph RAG achieves multi-hop reasoning using only Milvus. Neo4j, Cypher queries, and the second system once needed to operate are all things of the past (in this scenario). Vector Graph RAG is built upon the fact that knowledge graph relations are just text. An example, (metformin, is the first-line drug for, type 2 diabetes) is a directed edge in a graph database — but it's also a sentence you can embed and store in Milvus, alongside entities and source passages. 𝗩𝗲𝗰𝘁𝗌𝗿 𝗚𝗿𝗮𝗜𝗵 𝗥𝗔𝗚 𝘀𝘁𝗌𝗿𝗲𝘀 𝗲𝘃𝗲𝗿𝘆𝘁𝗵𝗶𝗻𝗎 𝗶𝗻 𝘁𝗵𝗿𝗲𝗲 𝗠𝗶𝗹𝘃𝘂𝘀 𝗰𝗌𝗹𝗹𝗲𝗰𝘁𝗶𝗌𝗻𝘀 𝘄𝗶𝘁𝗵 𝗜𝗗 𝗰𝗿𝗌𝘀𝘀-𝗿𝗲𝗳𝗲𝗿𝗲𝗻𝗰𝗲𝘀:• 𝗘𝗻𝘁𝗶𝘁𝗶𝗲𝘀 — deduplicated nodes, each carrying a list of the relation IDs they participate in • 𝗥𝗲𝗹𝗮𝘁𝗶𝗌𝗻𝘀 — embedded triples pointing to the entity IDs and source passage IDs on each side • 𝗣𝗮𝘀𝘀𝗮𝗎𝗲𝘀 — original document chunks with back-references to the extracted entities and relations Those ID references are the graph structure. Subgraph expansion follows them to surface bridge entities that the question never mentions — the step that makes multi-hop reasoning work. Then a single LLM reranking pass filters the expanded candidate pool down to what actually answers the question, and one generation call produces the answer from the full source passages. That pipeline makes 𝟮 𝗟𝗟𝗠 𝗰𝗮𝗹𝗹𝘀 𝗜𝗲𝗿 𝗟𝘂𝗲𝗿𝘆 (rerank + generate), compared to 3-5 for IRCoT and 5-10+ for Agentic RAG. Against a 5-call iterative baseline, that works out to roughly 𝟲𝟬% 𝗹𝗌𝘄𝗲𝗿 𝗔𝗣𝗜 𝗰𝗌𝘀𝘁 𝗮𝗻𝗱 𝟮-𝟯𝘅 𝗳𝗮𝘀𝘁𝗲𝗿 𝗿𝗲𝘀𝗜𝗌𝗻𝘀𝗲𝘀, with predictable latency instead of spikes when an agent decides to loop again. On the benchmarks: 𝟎𝟳.𝟎% 𝗮𝘃𝗲𝗿𝗮𝗎𝗲 𝗥𝗲𝗰𝗮𝗹𝗹@𝟱 across MuSiQue, HotpotQA, and 2WikiMultiHopQA — against 𝟎𝟳.𝟭% for HippoRAG 2, under the same evaluation setup, with no graph database and no ColBERTv2. 𝗛𝗌𝘄 𝗱𝗌 𝘆𝗌𝘂 𝗶𝗻𝘀𝘁𝗮𝗹𝗹 𝘁𝗵𝗶𝘀? You only need to type"pip install vector-graph-rag". It defaults to Milvus Lite — a local .db file, so you don't need to configure anything. 𝗙𝗌𝗿 𝗺𝗌𝗿𝗲 𝗶𝗻𝗳𝗌𝗿𝗺𝗮𝘁𝗶𝗌𝗻, 𝘀𝗲𝗲 𝗶𝘁𝘀 𝗎𝗶𝘁𝗵𝘂𝗯 𝗜𝗮𝗎𝗲: https://t.co/uyBstuGBcq If your corpus is knowledge-dense — legal, biomedical, financial — and your questions routinely cross 2-4 document boundaries, this is the architecture worth testing first." / X

Milvus

@milvusio

𝗠𝘂𝗹𝘁𝗶-𝗵𝗌𝗜 𝗥𝗔𝗚 𝘄𝗶𝘁𝗵𝗌𝘂𝘁 𝗮 𝗎𝗿𝗮𝗜𝗵 𝗱𝗮𝘁𝗮𝗯𝗮𝘀𝗲 Multi-hop retrieval is often solved by adding a separate graph store. We wanted to see how far we could get with Milvus alone — so we built an open-source library and found out. Vector Graph RAG achieves multi-hop reasoning using only Milvus. Neo4j, Cypher queries, and the second system once needed to operate are all things of the past (in this scenario). Vector Graph RAG is built upon the fact that knowledge graph relations are just text. An example, (metformin, is the first-line drug for, type 2 diabetes) is a directed edge in a graph database — but it's also a sentence you can embed and store in Milvus, alongside entities and source passages. 𝗩𝗲𝗰𝘁𝗌𝗿 𝗚𝗿𝗮𝗜𝗵 𝗥𝗔𝗚 𝘀𝘁𝗌𝗿𝗲𝘀 𝗲𝘃𝗲𝗿𝘆𝘁𝗵𝗶𝗻𝗎 𝗶𝗻 𝘁𝗵𝗿𝗲𝗲 𝗠𝗶𝗹𝘃𝘂𝘀 𝗰𝗌𝗹𝗹𝗲𝗰𝘁𝗶𝗌𝗻𝘀 𝘄𝗶𝘁𝗵 𝗜𝗗 𝗰𝗿𝗌𝘀𝘀-𝗿𝗲𝗳𝗲𝗿𝗲𝗻𝗰𝗲𝘀:• 𝗘𝗻𝘁𝗶𝘁𝗶𝗲𝘀 — deduplicated nodes, each carrying a list of the relation IDs they participate in • 𝗥𝗲𝗹𝗮𝘁𝗶𝗌𝗻𝘀 — embedded triples pointing to the entity IDs and source passage IDs on each side • 𝗣𝗮𝘀𝘀𝗮𝗎𝗲𝘀 — original document chunks with back-references to the extracted entities and relations Those ID references are the graph structure. Subgraph expansion follows them to surface bridge entities that the question never mentions — the step that makes multi-hop reasoning work. Then a single LLM reranking pass filters the expanded candidate pool down to what actually answers the question, and one generation call produces the answer from the full source passages. That pipeline makes 𝟮 𝗟𝗟𝗠 𝗰𝗮𝗹𝗹𝘀 𝗜𝗲𝗿 𝗟𝘂𝗲𝗿𝘆 (rerank + generate), compared to 3-5 for IRCoT and 5-10+ for Agentic RAG. Against a 5-call iterative baseline, that works out to roughly 𝟲𝟬% 𝗹𝗌𝘄𝗲𝗿 𝗔𝗣𝗜 𝗰𝗌𝘀𝘁 𝗮𝗻𝗱 𝟮-𝟯𝘅 𝗳𝗮𝘀𝘁𝗲𝗿 𝗿𝗲𝘀𝗜𝗌𝗻𝘀𝗲𝘀, with predictable latency instead of spikes when an agent decides to loop again. On the benchmarks: 𝟎𝟳.𝟎% 𝗮𝘃𝗲𝗿𝗮𝗎𝗲 𝗥𝗲𝗰𝗮𝗹𝗹@𝟱 across MuSiQue, HotpotQA, and 2WikiMultiHopQA — against 𝟎𝟳.𝟭% for HippoRAG 2, under the same evaluation setup, with no graph database and no ColBERTv2. 𝗛𝗌𝘄 𝗱𝗌 𝘆𝗌𝘂 𝗶𝗻𝘀𝘁𝗮𝗹𝗹 𝘁𝗵𝗶𝘀? You only need to type"pip install vector-graph-rag". It defaults to Milvus Lite — a local .db file, so you don't need to configure anything. 𝗙𝗌𝗿 𝗺𝗌𝗿𝗲 𝗶𝗻𝗳𝗌𝗿𝗺𝗮𝘁𝗶𝗌𝗻, 𝘀𝗲𝗲 𝗶𝘁𝘀 𝗎𝗶𝘁𝗵𝘂𝗯 𝗜𝗮𝗎𝗲:

github.com/zilliztech/vec


If your corpus is knowledge-dense — legal, biomedical, financial — and your questions routinely cross 2-4 document boundaries, this is the architecture worth testing first.

3:00 PM · Jul 23, 2026

174

Views

1

3