Milvus(@milvusio)

𝗬𝗌𝘂 𝗰𝗮𝗻 𝗿𝘂𝗻 𝟮𝟱 𝗺𝗶𝗹𝗹𝗶𝗌𝗻 𝘃𝗲𝗰𝘁𝗌𝗿𝘀 𝗶𝗻 𝗠𝗶𝗹𝘃𝘂𝘀 𝘂𝘀𝗶𝗻𝗎 𝘂𝗻𝗱𝗲𝗿 ...

8.5内容莚量
𝗬𝗌𝘂 𝗰𝗮𝗻 𝗿𝘂𝗻 𝟮𝟱 𝗺𝗶𝗹𝗹𝗶𝗌𝗻 𝘃𝗲𝗰𝘁𝗌𝗿𝘀 𝗶𝗻 𝗠𝗶𝗹𝘃𝘂𝘀 𝘂𝘀𝗶𝗻𝗎 𝘂𝗻𝗱𝗲𝗿 ...

TL;DR · AI 摘芁

Milvus 可圚单机 32GB 内存䞋运行 2500 侇 1280 绎囟像向量通过 FP16、mmap 和标量过滀技术实现。

栞心芁点

  • 䜿甚 FP16 可将每䞪向量绎床占甚内存从 4 字节减少到 2 字节。
  • mmap 技术允讞 Milvus 通过内存映射文件访问原始向量数据无需加蜜党郚到内存。
  • 标量过滀可将查询范囎猩小到几千䞪向量星著提升查询效率。

结构提纲

按章节快速跳蜬。

  1. 介绍甚户圚单机 32GB 内存䞋运行 2500 䞇囟像向量的挑战。

  2. 甚户尝试䜿甚 AI-SAQ 和 IVF_FLAT 玢匕䜆均未成功。

  3. 甚户最终选择 FLAT 玢匕结合 FP16、mmap 和标量过滀技术实现目标。

  4. FP16 将每䞪向量绎床占甚内存从 4 字节减少到 2 字节。

  5. mmap 技术允讞 Milvus 通过内存映射文件访问原始向量数据。

  6. 标量过滀可将查询范囎猩小到几千䞪向量星著提升查询效率。

思绎富囟

甚䞀匠囟看枅䞻题之闎的关系。

查看倧纲文本无障碍 / 无 JS 友奜
  • Milvus 内存䌘化方案
    • FP16 䌘化
      • 减少每䞪向量绎床占甚内存
    • mmap 技术
      • 通过内存映射文件访问原始向量数据
    • 标量过滀
      • 猩小查询范囎提升查询效率

金句 / Highlights

倌埗收藏䞎分享的关键句。

  • 䜿甚 FP16 可将每䞪向量绎床占甚内存从 4 字节减少到 2 字节减少原始向量数据占甚空闎的䞀半。

    — 第 3 段

    ⬇ 䞋蜜 PNG𝕏 分享到 X
  • mmap 技术允讞 Milvus 通过内存映射文件访问原始向量数据无需加蜜党郚到内存。

    — 第 3 段

    ⬇ 䞋蜜 PNG𝕏 分享到 X
  • 标量过滀可将查询范囎猩小到几千䞪向量星著提升查询效率。

    — 第 3 段

    ⬇ 䞋蜜 PNG𝕏 分享到 X
#Milvus#向量数据库#FP16#内存䌘化
打匀原文

Milvus on X: "𝗬𝗌𝘂 𝗰𝗮𝗻 𝗿𝘂𝗻 𝟮𝟱 𝗺𝗶𝗹𝗹𝗶𝗌𝗻 𝘃𝗲𝗰𝘁𝗌𝗿𝘀 𝗶𝗻 𝗠𝗶𝗹𝘃𝘂𝘀 𝘂𝘀𝗶𝗻𝗎 𝘂𝗻𝗱𝗲𝗿 𝟭𝗚𝗕 𝗌𝗳 𝗺𝗲𝗺𝗌𝗿𝘆. A user had 25M image vectors, each with 1280 dimensions, and only 32GB of memory available for Milvus on a single machine. The default FP32 sizing estimate https://t.co/HjjOXTSMCl" / X

Milvus

@milvusio

𝗬𝗌𝘂 𝗰𝗮𝗻 𝗿𝘂𝗻 𝟮𝟱 𝗺𝗶𝗹𝗹𝗶𝗌𝗻 𝘃𝗲𝗰𝘁𝗌𝗿𝘀 𝗶𝗻 𝗠𝗶𝗹𝘃𝘂𝘀 𝘂𝘀𝗶𝗻𝗎 𝘂𝗻𝗱𝗲𝗿 𝟭𝗚𝗕 𝗌𝗳 𝗺𝗲𝗺𝗌𝗿𝘆. A user had 25M image vectors, each with 1280 dimensions, and only 32GB of memory available for Milvus on a single machine. The default FP32 sizing estimate

tried more advanced indexes, but neither worked out: • 𝗔𝗜𝗊𝗔𝗀 looked right for constrained hardware, but the build path was too heavy for the machine. • 𝗜𝗩𝗙_𝗙𝗟𝗔𝗧 built successfully, but the collection load hung at 14% and never finished. After working with our developers, the user switched to 𝗙𝗟𝗔𝗧, the simplest index in Milvus. FLAT avoided extra ANN structures and build/load complexity, while Milvus provided the pieces that made the setup practical: • 𝗙𝗣𝟭𝟲 storage cut each vector dimension from 4 bytes to 2 bytes, reducing raw vector data by half. • 𝗺𝗺𝗮𝗜 let Milvus access raw vector data through memory-mapped files instead of loading it all into process memory. • 𝗊𝗰𝗮𝗹𝗮𝗿 𝗳𝗶𝗹𝘁𝗲𝗿𝗶𝗻𝗎 narrowed each query first using fields like dataid and classid, so Milvus compared only a few thousand vectors instead of 25 million. 𝗧𝗵𝗲 𝗿𝗲𝘀𝘂𝗹𝘁: 𝗮𝗿𝗌𝘂𝗻𝗱 𝟲𝟬𝟬𝗠𝗕 𝗌𝗳 𝗿𝗲𝘀𝗶𝗱𝗲𝗻𝘁 𝗺𝗲𝗺𝗌𝗿𝘆 𝗮𝗻𝗱 𝘄𝗮𝗿𝗺 𝗟𝘂𝗲𝗿𝗶𝗲𝘀 𝘂𝗻𝗱𝗲𝗿 𝟭𝟬𝟬𝗺𝘀. When the real search space is much smaller than the full collection, as in multi-tenant RAG, labeled image search, or e-commerce search, FLAT + FP16 + mmap can be a practical option. Full breakdown in the blog:

milvus.io/blog/25-millio


5:58 PM · Jun 5, 2026

983

Views

2

1

3

13

6

16