We talked to the authors of ð¥ð®ðð¶ðð€, the people affected by Google's TurboQuant paper. ...

TL;DR · AI æèŠ
æ žå¿èŠç¹
ðð¶ðð² ððµð¶ð»ðŽð ðððð°ðž ðð¶ððµ ðð: ⢠ð©ð²ð°ððŒð¿ ðŸðð®ð»ðð¶ðð®ðð¶ðŒð» ðµð®ð ðµð¶ð ð¶ðð ððµð²ðŒð¿ð²ðð¶ð°ð®ð¹ ð°ð²ð¶ð¹ð¶ð»ðŽ. RaBitQ is mathematically proven https://t.co/n426B4ZY6w" / X
We talked to the authors of ð¥ð®ðð¶ðð€, the people affected by Google's TurboQuant paper. ðð¶ðð² ððµð¶ð»ðŽð ðððð°ðž ðð¶ððµ ðð: ⢠ð©ð²ð°ððŒð¿ ðŸðð®ð»ðð¶ðð®ðð¶ðŒð» ðµð®ð ðµð¶ð ð¶ðð ððµð²ðŒð¿ð²ðð¶ð°ð®ð¹ ð°ð²ð¶ð¹ð¶ð»ðŽ. RaBitQ is mathematically proven asymptotically optimal. The remaining gains are on the engineering side: hardware, data distribution, latency. ⢠ððŒðºðœð¿ð²ððð¶ðŒð» ððŒð»'ð ððµð¿ð¶ð»ðž ðððŒð¿ð®ðŽð² ð±ð²ðºð®ð»ð±. ðð ðºð¶ðŽðµð ðŽð¿ðŒð ð¶ð. Smaller vectors mean larger models run on smaller devices, which creates new workloads instead of replacing old ones. ⢠ðŠð¶ð»ð°ð² ð¥ð®ðð¶ðð€, ðºð®ððµð²ðºð®ðð¶ð°ð®ð¹ ðð²ð°ððŒð¿ ðŸðð®ð»ðð¶ðð®ðð¶ðŒð» ð¶ð ððµð¿ð²ð² ððð²ðœð. Random rotation (a form of Johnson-Lindenstrauss transformation) to spread information evenly across dimensions, grid construction, then quantization. ⢠ð©ð²ð°ððŒð¿ ðŸðð®ð»ðð¶ðð®ðð¶ðŒð» ð¶ð ð¶ð»ð³ð¹ðð²ð»ð°ð¶ð»ðŽ ðð¿ð®ð»ðð³ðŒð¿ðºð²ð¿ ð¶ð»ð³ð²ð¿ð²ð»ð°ð². On the surface, KV cache compression and ANN vector compression are different problems. Mathematically, they share most of the same logic. ⢠ðð© ð°ð®ð°ðµð² ð¶ð ð°ðµð²ð®ðœ ðððŒð¿ð®ðŽð² ðð¿ð®ð±ð²ð± ð³ðŒð¿ ð²ð ðœð²ð»ðð¶ðð² ð°ðŒðºðœððð². Quantization makes that trade more favorable on both sides. ⢠ðªð®ð»ð ððŒ ðð²ð² ð²ð»ðŽð¶ð»ð²ð²ð¿ð²ð± ð¥ð®ðð¶ðð€ ð¶ð» ðœð¿ðŒð±ðð°ðð¶ðŒð»? Try ðð©ð_ð¥ðððð§ð€ ð¶ð» ð ð¶ð¹ððð ð®.ð². ð¡ðŒðð²: The views expressed are those of the interviewees and do not represent Zilliz. Views belong to ðð¶ð®ð»ðð®ð»ðŽ ðð®ðŒ (RaBitQ first author), ððµð²ð»ðŽ ððŒð»ðŽ (RaBitQ co-author), and ðð¶ ðð¶ð (Zilliz Engineering). ððð¹ð¹ ð°ðŒð»ðð²ð¿ðð®ðð¶ðŒð»: milvus.io/blog/interview