Milvus(@milvusio)

We talked to the authors of 𝗥𝗮𝗕𝗶𝘁𝗀, the people affected by Google's TurboQuant paper. ...

5.0内容莚量
We talked to the authors of 𝗥𝗮𝗕𝗶𝘁𝗀, the people affected by Google's TurboQuant paper. 

...

TL;DR · AI 摘芁

栞心芁点

𝗙𝗶𝘃𝗲 𝘁𝗵𝗶𝗻𝗎𝘀 𝘀𝘁𝘂𝗰𝗞 𝘄𝗶𝘁𝗵 𝘂𝘀: • 𝗩𝗲𝗰𝘁𝗌𝗿 𝗟𝘂𝗮𝗻𝘁𝗶𝘇𝗮𝘁𝗶𝗌𝗻 𝗵𝗮𝘀 𝗵𝗶𝘁 𝗶𝘁𝘀 𝘁𝗵𝗲𝗌𝗿𝗲𝘁𝗶𝗰𝗮𝗹 𝗰𝗲𝗶𝗹𝗶𝗻𝗎. RaBitQ is mathematically proven https://t.co/n426B4ZY6w" / X

We talked to the authors of 𝗥𝗮𝗕𝗶𝘁𝗀, the people affected by Google's TurboQuant paper. 𝗙𝗶𝘃𝗲 𝘁𝗵𝗶𝗻𝗎𝘀 𝘀𝘁𝘂𝗰𝗞 𝘄𝗶𝘁𝗵 𝘂𝘀: • 𝗩𝗲𝗰𝘁𝗌𝗿 𝗟𝘂𝗮𝗻𝘁𝗶𝘇𝗮𝘁𝗶𝗌𝗻 𝗵𝗮𝘀 𝗵𝗶𝘁 𝗶𝘁𝘀 𝘁𝗵𝗲𝗌𝗿𝗲𝘁𝗶𝗰𝗮𝗹 𝗰𝗲𝗶𝗹𝗶𝗻𝗎. RaBitQ is mathematically proven asymptotically optimal. The remaining gains are on the engineering side: hardware, data distribution, latency. • 𝗖𝗌𝗺𝗜𝗿𝗲𝘀𝘀𝗶𝗌𝗻 𝘄𝗌𝗻'𝘁 𝘀𝗵𝗿𝗶𝗻𝗞 𝘀𝘁𝗌𝗿𝗮𝗎𝗲 𝗱𝗲𝗺𝗮𝗻𝗱. 𝗜𝘁 𝗺𝗶𝗎𝗵𝘁 𝗎𝗿𝗌𝘄 𝗶𝘁. Smaller vectors mean larger models run on smaller devices, which creates new workloads instead of replacing old ones. • 𝗊𝗶𝗻𝗰𝗲 𝗥𝗮𝗕𝗶𝘁𝗀, 𝗺𝗮𝘁𝗵𝗲𝗺𝗮𝘁𝗶𝗰𝗮𝗹 𝘃𝗲𝗰𝘁𝗌𝗿 𝗟𝘂𝗮𝗻𝘁𝗶𝘇𝗮𝘁𝗶𝗌𝗻 𝗶𝘀 𝘁𝗵𝗿𝗲𝗲 𝘀𝘁𝗲𝗜𝘀. Random rotation (a form of Johnson-Lindenstrauss transformation) to spread information evenly across dimensions, grid construction, then quantization. • 𝗩𝗲𝗰𝘁𝗌𝗿 𝗟𝘂𝗮𝗻𝘁𝗶𝘇𝗮𝘁𝗶𝗌𝗻 𝗶𝘀 𝗶𝗻𝗳𝗹𝘂𝗲𝗻𝗰𝗶𝗻𝗎 𝘁𝗿𝗮𝗻𝘀𝗳𝗌𝗿𝗺𝗲𝗿 𝗶𝗻𝗳𝗲𝗿𝗲𝗻𝗰𝗲. On the surface, KV cache compression and ANN vector compression are different problems. Mathematically, they share most of the same logic. • 𝗞𝗩 𝗰𝗮𝗰𝗵𝗲 𝗶𝘀 𝗰𝗵𝗲𝗮𝗜 𝘀𝘁𝗌𝗿𝗮𝗎𝗲 𝘁𝗿𝗮𝗱𝗲𝗱 𝗳𝗌𝗿 𝗲𝘅𝗜𝗲𝗻𝘀𝗶𝘃𝗲 𝗰𝗌𝗺𝗜𝘂𝘁𝗲. Quantization makes that trade more favorable on both sides. • 𝗪𝗮𝗻𝘁 𝘁𝗌 𝘀𝗲𝗲 𝗲𝗻𝗎𝗶𝗻𝗲𝗲𝗿𝗲𝗱 𝗥𝗮𝗕𝗶𝘁𝗀 𝗶𝗻 𝗜𝗿𝗌𝗱𝘂𝗰𝘁𝗶𝗌𝗻? Try 𝗜𝗩𝗙_𝗥𝗔𝗕𝗜𝗧𝗀 𝗶𝗻 𝗠𝗶𝗹𝘃𝘂𝘀 𝟮.𝟲. 𝗡𝗌𝘁𝗲: The views expressed are those of the interviewees and do not represent Zilliz. Views belong to 𝗝𝗶𝗮𝗻𝘆𝗮𝗻𝗎 𝗚𝗮𝗌 (RaBitQ first author), 𝗖𝗵𝗲𝗻𝗎 𝗟𝗌𝗻𝗎 (RaBitQ co-author), and 𝗟𝗶 𝗟𝗶𝘂 (Zilliz Engineering). 𝗙𝘂𝗹𝗹 𝗰𝗌𝗻𝘃𝗲𝗿𝘀𝗮𝘁𝗶𝗌𝗻: milvus.io/blog/interview

Image 1: Image
Image 1: Image