Google’s TurboQuant Compression May Support Faster Inference, Same Accuracy on Less Capable Hardware
TL;DR · AI 摘要
本文介绍了Google推出的TurboQuant压缩技术,旨在通过优化KV Cache实现大模型推理加速,在保持精度不变的前提下降低对硬件算力的要求。但提供文本缺失核心机制与实验数据,技术细节不足。
核心要点
- Google提出TurboQuant压缩算法,针对大模型KV Cache进行优化。
- 该技术宣称可在低算力硬件上实现推理加速且保持模型精度不变。
- 原文缺失具体架构设计、量化策略与基准测试数据,工程参考价值有限。
Google’s TurboQuant Compression May Support Faster Inference, Same Accuracy on Less Capable Hardware - InfoQ
[BT](http://www.infoq.com/int/bt/ "bt")
InfoQ Software Architects' Newsletter
A monthly overview of things you need to know as an architect or aspiring architect.
Enter your e-mail address
Select your country - [x] I consent to InfoQ.com handling my data as explained in this Privacy Notice.
Close
Live Webinar and Q&A: Designing Data Layers for Agentic AI: Patterns for State, Memory, and Coordination at Scale (May 12, 2026)Save Your Seat
Close
Toggle Navigation
Facilitating the Spread of Knowledge and Innovation in Professional Software Development
English edition
[Write for InfoQ](http://www.infoq.com/write-for-infoq/ "Write for InfoQ")
Search
Unlock the full InfoQ experience
Unlock the full InfoQ experience by logging in! Stay updated with your favorite authors and topics, engage with content, and download exclusive resources.
or
Don't have an InfoQ account?
- Stay updated on topics and peers that matter to youReceive instant alerts on the latest insights and trends.
- Quickly access free resources for continuous learningMinibooks, videos with transcripts, and training materials.
- Save articles and read at anytimeBookmark articles to read whenever youre ready.
NewsArticlesPresentationsPodcastsGuides
Topics
[Development](http://www.infoq.com/development/ "Development")
- [Java](http://www.infoq.com/java/ "Java")
- [Kotlin](http://www.infoq.com/kotlin/ "Kotlin")
- [.Net](http://www.infoq.com/dotnet/ ".Net")
- [C#](http://www.infoq.com/c_sharp/ "C#")
- [Swift](http://www.infoq.com/swift/ "Swift")
- [Go](http://www.infoq.com/golang/ "Go")
- [Rust](http://www.infoq.com/rust/ "Rust")
- [JavaScript](http://www.infoq.com/javascript/ "JavaScript")
Featured in Development
Agent workflows make transport a first-order concern. Multi-turn, tool-heavy loops amplify overhead that is negligible in single-turn LLM use. Stateful continuation cuts overhead dramatically. Caching context server-side can reduce client-sent data by 80%+ and improve execution time by 15–29% .

All in developmentFollow Topic
[Architecture & Design](http://www.infoq.com/architecture-design/ "Architecture & Design")
- [Architecture](http://www.infoq.com/architecture/ "Architecture")
- [Enterprise Architecture](http://www.infoq.com/enterprise-architecture/ "Enterprise Architecture")
- [Scalability/Performance](http://www.infoq.com/performance-scalability/ "Scalability/Performance")
- [Design](http://www.infoq.com/design/ "Design")
- [Case Studies](http://www.infoq.com/Case_Study/ "Case Studies")
- [Microservices](http://www.infoq.com/microservices/ "Microservices")
- [Service Mesh](http://www.infoq.com/servicemesh/ "Service Mesh")
- [Patterns](http://www.infoq.com/DesignPattern/ "Patterns")
- [Security](http://www.infoq.com/Security/ "Security")
Featured in Architecture & Design
Randy Shoup discusses the "Velocity Initiative," a transformation that doubled engineering productivity and modernized eBay’s DORA metrics. He shares the technical playbook used to scale 4,500 services while explaining why even elite engineering execution can’t save a company hampered by waterfall planning, risk aversion, and a "pathological" culture of fear.

All in architecture-designFollow Topic
[AI Infrastructure](http://www.infoq.com/ai-ml-data-eng/ "AI Infrastructure")
- [Big Data](http://www.infoq.com/bigdata/ "Big Data")
- [Machine Learning](http://www.infoq.com/machinelearning/ "Machine Learning")
- [NoSQL](http://www.infoq.com/nosql/ "NoSQL")
- [Database](http://www.infoq.com/database/ "Database")
- [Data Analytics](http://www.infoq.com/data-analytics/ "Data Analytics")
- [Streaming](http://www.infoq.com/streaming/ "Streaming")
Featured in AI, ML & Data Engineering
Mariia Bulycheva discusses the transition from classic deep learning to GNNs for Zalando's landing page. She explains the complexities of converting user logs into heterogeneous graphs, the "message passing" training process, and the technical pitfalls of graph data leakage. She shares how a hybrid architecture solved inference latency, delivering contextual embeddings to a downstream model.

All in ai-ml-data-engFollow Topic
[Culture & Methods](http://www.infoq.com/culture-methods/ "Culture & Methods")
- [Agile](http://www.infoq.com/agile/ "Agile")
- [Diversity](http://www.infoq.com/diversity/ "Diversity")
- [Leadership](http://www.infoq.com/leadership/ "Leadership")
- [Lean/Kanban](http://www.infoq.com/lean/ "Lean/Kanban")
- [Personal Growth](http://www.infoq.com/personal-growth/ "Personal Growth")
- [Scrum](http://www.infoq.com/scrum/ "Scrum")
- [Sociocracy](http://www.infoq.com/sociocracy/ "Sociocracy")
- [Software Craftmanship](http://www.infoq.com/software_craftsmanship/ "Software Craftmanship")
- [Team Collaboration](http://www.infoq.com/team-collaboration/ "Team Collaboration")
- [Testing](http://www.infoq.com/testing/ "Testing")
- [UX](http://www.infoq.com/ux/ "UX")
Featured in Culture & Methods
Celine Pypaert discusses the ubiquitous nature of open-source software and shares a blueprint for securing modern applications. She explains how to prioritize high-risk vulnerabilities using exploitability data, the role of Software Bill of Materials (SBOM), and the importance of bridging the gap between DevOps and Security through clear accountability and automated governance.

All in culture-methodsFollow Topic
- [Infrastructure](http://www.infoq.com/infrastructure/ "Infrastructure")
- [Continuous Delivery](http://www.infoq.com/continuous_delivery/ "Continuous Delivery")
- [Automation](http://www.infoq.com/automation/ "Automation")
- [Containers](http://www.infoq.com/containers/ "Containers")
- [Cloud](http://www.infoq.com/cloud-computing/ "Cloud")
- [Observability](http://www.infoq.com/observability/ "Observability")
Featured in DevOps
Docker Extensions boost developer speed but create a "visibility gap" by isolating telemetry. To meet enterprise needs, extensions must act as bridges to centralized platforms. This article details how to use OpenTelemetry, policy-as-code, and encryption to build secure pipelines. Learn to balance developer productivity with the governance required for scalable, compliant observability.

All in devopsFollow Topic
[Events](https://events.infoq.com/ "Events")
Helpful links
- [About InfoQ](http://www.infoq.com/about-infoq "About InfoQ")
- [InfoQ Editors](http://www.infoq.com/infoq-editors "InfoQ Editors")
- [Write for InfoQ](http://www.infoq.com/write-for-infoq "Write for InfoQ")
- [About C4Media](https://c4media.com/ "About C4Media")
- [Diversity](https://c4media.com/diversity "Diversity")
Choose your language

[InfoQ Homepage](http://www.infoq.com/ "InfoQ Homepage")[News](http://www.infoq.com/news "News")Google’s TurboQuant Compression May Support Faster Inference, Same Accuracy on Less Capable Hardware
[AI, ML & Data Engineering](http://www.infoq.com/ai-ml-data-eng/ "AI, ML & Data Engineering")
Shipping Faster, Breaking More: Rethinking Delivery Systems in the Age of AI (Webinar May 28th)
Google’s TurboQuant Compression May Support Faster Inference, Same Accuracy on Less Capable Hardware
Apr 15, 2026 3 min read
by
- [](http://www.infoq.com/profile/Bruno-Couriol/)Bruno Couriol
Follow Application Consultant
#### Write for InfoQ
Feed your curiosity.Help 550k+ global
senior developers
each month stay ahead.Get in touch
Log in to listen to this article
Loading audio
Your browser does not support the audio element.
0:00 0:00
Normal 1.25x 1.5x
Like
Google Research unveiled TurboQuant, a novel quantization algorithm that compresses large language models’ Key-Value caches by up to 6x. With 3.5-bit compression, near-zero accuracy loss, and no retraining needed, it allows developers to run massive context windows on significantly more modest hardware than previously required. Early community benchmarks confirm significant efficiency gains.
While the rationale for quantization may seem logical, the difficulty is, given a number of encoding bits, to maintain the accuracy of inference-relevant computations (e.g., inner products, cosine similarity, distances) with compressed data.
The research team claims that TurboQuant can compress the KV cache down to 3.5 bits per value with near-zero accuracy loss. On standard benchmarks like LongBench and Needle in a Haystack, a 3.5-bit TurboQuant implementation matched the performance of full 16-bit precision across Gemma and Mistral models.
TurboQuant uses a two-step approach. First, data vectors are rotated (randomized Hadamard transform). This keeps key Euclidean properties (e.g., distance) while spreading out the values, removing the outlier-heavy coordinate distribution that makes low-bit quantization difficult. Post-transform, the vector coordinates follow a beta distribution that is more amenable to compression with low distortion. Second, a decade-old technique, the Quantized Johnson-Lindenstrauss (QJL) transform, is applied to remove the bias created by the first step. Post-QJL, the paper argues that inner products between quantized vectors are unbiased, computationally efficient, accurate estimators of the unquantized vectors, with the resulting effect of maintaining inference accuracy.
Early community analysis seems to confirm significant gains, albeit more modest than those reported in the paper. The Two Minute Papers analysis suggests more realistic, “real-world” improvements of 30-40% in memory reduction and processing speed:
Based on the results, we cannot conclude that every AI machine suddenly needs 6 times less RAM. No. That is a bit idealistic and only true for some corner cases. You know when you see an official benchmark of a phone battery or electric car mileage with somewhat idealized conditions? It is a bit like that.
So careful with the media hype. […] We wait for more data and analyze experiments here, to get the highest quality information.
But it’s still good. Really good! It helps most people who run AI systems with very long contexts. When you chuck in a huge PDF document, or a movie, or a huge codebase for an AI to analyze. Yes, you will be able to do that cheaper, with meaningfully less memory. Often a few gigabytes less. And I think that is absolutely amazing news.
A fundamental optimization in LLM inference is the caching of computations that are required repeatedly. This is particularly critical during autoregressive generation, where each newly generated token uses data already computed during the generation of all previous tokens. By caching these Key and Value tensors (the KV cache), the system avoids redundant, computationally expensive passes over the entire sequence history.
However, the efficiency gains offered by caching come with a significant memory cost that grows linearly with the token sequence length. For LLMs designed with long context windows, the massive VRAM footprint of the cache eventually outweighs the memory required for the model weights themselves.
For example, according to Darshan Fofadiya, AI researcher at Amazon,running a Llama 70B model with a 1M-token context window may require approximately 328GB of VRAM just for the KV cache. When compared to the 140 GB required to hold the 70B model weights in BF16, the cache becomes the primary barrier to deployment, forcing engineers into costly multi-GPU configurations. Compressed down from 16 to 3.5 bits, the cache then requires 72 GB and fits on a single H100 (80 GB HBM).
During the inference decoding phase, certain input tokens in a prompt produce KV vectors with magnitudes in the hundreds or thousands, while the majority of other tokens have values close to 0-1. In LLaMA-2-7B, for instance, the top 1% of KV cache values may have magnitudes that are 10-100x larger than the median value. This massive distribution skew makes linear 4-bit quantization impossible without specialized techniques, as the outliers stretch the quantization grid and crush the precision of normal tokens.
Generative inference with LLMs is, for relatively small batch sizes, memory-bound. With memory speeds growing slower than compute speeds, reducing the memory bottleneck (the so-called memory wall) is key to efficient inference. For short contexts, weight matrices are the dominant contributor to memory consumption. For long contexts, the KV cache becomes the main contributor. Quantization techniques for both model weights and the KV cache are thus instrumental in speeding up inference and constitute a major research topic.
About the Author
[](http://www.infoq.com/profile/Bruno-Couriol/)
#### Bruno Couriol
MSc in Telecommunications. BSc in Mathematics.
Show more Show less
#### This content is in the AI, ML & Data Engineering topic
Follow Topic
##### Related Topics:
Followers: 4084
Follow Topic
Followers: 5862
Follow Topic
Followers: 9
Follow Topic
Followers: 0
Follow Topic
Followers: 279
Follow Topic
Followers: 137
Follow Topic
* #### Popular in AI, ML & Data Engineering
- ##### Anthropic Releases Claude Mythos Preview with Cybersecurity Capabilities but Withholds Public Access
- ##### Building Hierarchical Agentic RAG Systems: Multi-Modal Reasoning with Autonomous Error Recovery
* #### Related Sponsors
- #### Related Sponsor
Confidently test, evaluate, and red-team your LLM apps with Promptfoo — catch regressions, benchmark models, and ship high-quality AI features faster; start testing your prompts today. [Learn More](http://www.infoq.com/url/f/0ed8a8f2-ad41-400e-b24f-e10459b3993d/).
Related Content
Apr 14, 2026
Apr 13, 2026
Apr 13, 2026
Apr 07, 2026
Mar 27, 2026
Mar 26, 2026
Mar 24, 2026
Mar 23, 2026
Mar 20, 2026
Related Sponsors
- #### Inside MCP: A Protocol for AI Integration
The Model Context Protocol (MCP) defines a standard way for AI systems to interact with tools, data, and services. This article explains MCP’s architecture—hosts, clients, and servers—and how it enables structured, secure integrations between AI models and external systems.
- #### Harder, Better, Prompter, Stronger: AI system prompt hardening
System prompts define how LLM applications behave—but they are vulnerable to manipulation. This article explores prompt hardening techniques such as instruction shielding, syntax reinforcement, and layered prompting to defend AI systems against prompt injection and override attacks.
- Sponsored by

Related Content
- ##### Stripe Engineers Deploy Minions, Autonomous Agents Producing Thousands of Pull Requests Weekly
Mar 20, 2026
Mar 19, 2026
Apr 02, 2026 
Feb 13, 2026 
Feb 09, 2026 
- ##### Article Series: AI-Assisted Development: Real World Patterns, Pitfalls, and Production Readiness
Jan 21, 2026 
**The InfoQ** Newsletter
A round-up of last week’s content on InfoQ sent out every Tuesday. Join a community of over 250,000 senior developers. View an example
Enter your e-mail address
Select your country - [x] I consent to InfoQ.com handling my data as explained in this Privacy Notice.
- ##### [Claude Code Used to Find Remotely Exploitable Linux Kernel Vulnerability Hidden for 23 Years](http://www.infoq.com/news/2026/04/claude-code-linux-vulnerability/ "Claude Code Used to Find Remotely Exploitable Linux Kernel Vulnerability Hidden for 23 Years")
- ##### [Cloudflare Introduces EmDash: TypeScript CMS Positioned as WordPress Successor](http://www.infoq.com/news/2026/04/cloudflare-emdash-wordpress/ "Cloudflare Introduces EmDash: TypeScript CMS Positioned as WordPress Successor")
- ##### [Stateful Continuation for AI Agents: Why Transport Layers Now Matter](http://www.infoq.com/articles/ai-agent-transport-layer/ "Stateful Continuation for AI Agents: Why Transport Layers Now Matter")
- ##### [Zendesk Says AI Makes Code Abundant, Shifting the Bottleneck to “Absorption Capacity”](http://www.infoq.com/news/2026/04/zendesk-absorption-capacity/ "Zendesk Says AI Makes Code Abundant, Shifting the Bottleneck to “Absorption Capacity”")
- ##### [Platform Engineering: Lessons from the Rise and Fall of eBay Velocity](http://www.infoq.com/presentations/platform-engineering-lessons/ "Platform Engineering: Lessons from the Rise and Fall of eBay Velocity")
- ##### [Lyft Scales Global Localization Using AI and Human-in-the-Loop Review](http://www.infoq.com/news/2026/04/lyft-ai-localization-pipeline/ "Lyft Scales Global Localization Using AI and Human-in-the-Loop Review")
- ##### [Empower Your Developers: How Open Source Dependencies Risk Management Can Unlock Innovation](http://www.infoq.com/presentations/open-source-dependencies/ "Empower Your Developers: How Open Source Dependencies Risk Management Can Unlock Innovation")
- ##### [Tiger Teams, Evals and Agents: The New AI Engineering Playbook](http://www.infoq.com/podcasts/tiger-teams-evals-agents/ "Tiger Teams, Evals and Agents: The New AI Engineering Playbook")
- ##### [Developing Your Leadership Skills toward Principal Engineering](http://www.infoq.com/news/2026/04/leadership-skills/ "Developing Your Leadership Skills toward Principal Engineering")
- ##### [Google’s TurboQuant Compression May Support Faster Inference, Same Accuracy on Less Capable Hardware](http://www.infoq.com/news/2026/04/turboquant-compression-kv-cache/ "Google’s TurboQuant Compression May Support Faster Inference, Same Accuracy on Less Capable Hardware")
- ##### [Anthropic Paper Examines Behavioral Impact of Emotion-Like Mechanisms in LLMs](http://www.infoq.com/news/2026/04/anthropic-paper-llms/ "Anthropic Paper Examines Behavioral Impact of Emotion-Like Mechanisms in LLMs")
- ##### [Reimagining Platform Engagement with Graph Neural Networks](http://www.infoq.com/presentations/graph-neural-networks/ "Reimagining Platform Engagement with Graph Neural Networks")
- ##### [OpenTelemetry Declarative Configuration Reaches Stability Milestone](http://www.infoq.com/news/2026/04/opentelemetry-declarative-config/ "OpenTelemetry Declarative Configuration Reaches Stability Milestone")
- ##### [New Rowhammer Attacks on NVIDIA GPUs Enable Full System Takeover](http://www.infoq.com/news/2026/04/rowhammer-attacks-nvidia/ "New Rowhammer Attacks on NVIDIA GPUs Enable Full System Takeover")
- ##### [Beyond One-Click: Designing an Enterprise-Grade Observability Extension for Docker](http://www.infoq.com/articles/enterprise-grade-observability-extension-docker/ "Beyond One-Click: Designing an Enterprise-Grade Observability Extension for Docker")
**The InfoQ** Newsletter
A round-up of last week’s content on InfoQ sent out every Tuesday. Join a community of over 250,000 senior developers. View an example
- Get a quick overview of content published on a variety of innovator and early adopter technologies
- Learn what you don’t know that you don’t know
- Stay up to date with the latest information from the topics you are interested in
Enter your e-mail address
Select your country - [x] I consent to InfoQ.com handling my data as explained in this Privacy Notice.
#### Events
April 15 | May 7 | June 10, 2026
- ##### QCon AI Boston
June 1-2, 2026
- ##### QCon San Francisco
November 16-20, 2026
#### Follow us on
Youtube 232K FollowersLinkedin 26K FollowersRSS 19K ReadersX 57.1k FollowersFacebook 21K LikesBluesky NewAlexa New
#### Stay in the know
The InfoQ PodcastEngineering Culture PodcastThe Software Architects' Newsletter
General Feedback [[email protected]](mailto:[email protected]) Advertising [[email protected]](mailto:[email protected]) Editorial [[email protected]](mailto:[email protected]) Marketing [[email protected]](mailto:[email protected])
InfoQ.com and all content copyright © 2006-2026 C4Media Inc.
Privacy Notice, Terms And Conditions, Cookie Policy
Close
[BT](http://www.infoq.com/int/bt/ "bt")