InfoQ

Google’s TurboQuant Compression May Support Faster Inference, Same Accuracy on Less Capable Hardware

3.5内容质量
Google’s TurboQuant Compression May Support Faster Inference, Same Accuracy on Less Capable Hardware

TL;DR · AI 摘要

本文介绍了Google推出的TurboQuant压缩技术,旨在通过优化KV Cache实现大模型推理加速,在保持精度不变的前提下降低对硬件算力的要求。但提供文本缺失核心机制与实验数据,技术细节不足。

核心要点

  • Google提出TurboQuant压缩算法,针对大模型KV Cache进行优化。
  • 该技术宣称可在低算力硬件上实现推理加速且保持模型精度不变。
  • 原文缺失具体架构设计、量化策略与基准测试数据,工程参考价值有限。
#大模型推理#KV Cache#模型量化#Google#AI基础设施
打开原文

Google’s TurboQuant Compression May Support Faster Inference, Same Accuracy on Less Capable Hardware - InfoQ

[BT](http://www.infoq.com/int/bt/ "bt")

InfoQ Software Architects' Newsletter

A monthly overview of things you need to know as an architect or aspiring architect.

View an example

Enter your e-mail address

Select your country - [x] I consent to InfoQ.com handling my data as explained in this Privacy Notice.

We protect your privacy.

Close

Live Webinar and Q&A: Designing Data Layers for Agentic AI: Patterns for State, Memory, and Coordination at Scale (May 12, 2026)Save Your Seat

Close

Toggle Navigation

Facilitating the Spread of Knowledge and Innovation in Professional Software Development

English edition

[Write for InfoQ](http://www.infoq.com/write-for-infoq/ "Write for InfoQ")

Search

RegisterSign in

Unlock the full InfoQ experience

Unlock the full InfoQ experience by logging in! Stay updated with your favorite authors and topics, engage with content, and download exclusive resources.

Log In

or

Don't have an InfoQ account?

Register

  • Stay updated on topics and peers that matter to youReceive instant alerts on the latest insights and trends.
  • Quickly access free resources for continuous learningMinibooks, videos with transcripts, and training materials.
  • Save articles and read at anytimeBookmark articles to read whenever youre ready.

Logo - Back to homepage

NewsArticlesPresentationsPodcastsGuides

Topics

[Development](http://www.infoq.com/development/ "Development")

  • [Java](http://www.infoq.com/java/ "Java")
  • [Kotlin](http://www.infoq.com/kotlin/ "Kotlin")
  • [.Net](http://www.infoq.com/dotnet/ ".Net")
  • [C#](http://www.infoq.com/c_sharp/ "C#")
  • [Swift](http://www.infoq.com/swift/ "Swift")
  • [Go](http://www.infoq.com/golang/ "Go")
  • [Rust](http://www.infoq.com/rust/ "Rust")
  • [JavaScript](http://www.infoq.com/javascript/ "JavaScript")

Featured in Development

Agent workflows make transport a first-order concern. Multi-turn, tool-heavy loops amplify overhead that is negligible in single-turn LLM use. Stateful continuation cuts overhead dramatically. Caching context server-side can reduce client-sent data by 80%+ and improve execution time by 15–29% .

![Image 1: Stateful Continuation for AI Agents: Why Transport Layers Now Matter/articles/ai-agent-transport-layer/en/smallimage/ai-agent-transport-layer-thumbnail-1775031603285.jpg)](http://www.infoq.com/articles/ai-agent-transport-layer)

All in developmentFollow Topic

[Architecture & Design](http://www.infoq.com/architecture-design/ "Architecture & Design")

  • [Architecture](http://www.infoq.com/architecture/ "Architecture")
  • [Enterprise Architecture](http://www.infoq.com/enterprise-architecture/ "Enterprise Architecture")
  • [Scalability/Performance](http://www.infoq.com/performance-scalability/ "Scalability/Performance")
  • [Design](http://www.infoq.com/design/ "Design")
  • [Case Studies](http://www.infoq.com/Case_Study/ "Case Studies")
  • [Microservices](http://www.infoq.com/microservices/ "Microservices")
  • [Service Mesh](http://www.infoq.com/servicemesh/ "Service Mesh")
  • [Patterns](http://www.infoq.com/DesignPattern/ "Patterns")
  • [Security](http://www.infoq.com/Security/ "Security")

Featured in Architecture & Design

Randy Shoup discusses the "Velocity Initiative," a transformation that doubled engineering productivity and modernized eBay’s DORA metrics. He shares the technical playbook used to scale 4,500 services while explaining why even elite engineering execution can’t save a company hampered by waterfall planning, risk aversion, and a "pathological" culture of fear.

![Image 2: Platform Engineering: Lessons from the Rise and Fall of eBay Velocity/presentations/platform-engineering-lessons/en/smallimage/randy-shoup-thumbnail-1775637120944.jpg)](http://www.infoq.com/presentations/platform-engineering-lessons)

All in architecture-designFollow Topic

[AI Infrastructure](http://www.infoq.com/ai-ml-data-eng/ "AI Infrastructure")

  • [Big Data](http://www.infoq.com/bigdata/ "Big Data")
  • [Machine Learning](http://www.infoq.com/machinelearning/ "Machine Learning")
  • [NoSQL](http://www.infoq.com/nosql/ "NoSQL")
  • [Database](http://www.infoq.com/database/ "Database")
  • [Data Analytics](http://www.infoq.com/data-analytics/ "Data Analytics")
  • [Streaming](http://www.infoq.com/streaming/ "Streaming")

Featured in AI, ML & Data Engineering

Mariia Bulycheva discusses the transition from classic deep learning to GNNs for Zalando's landing page. She explains the complexities of converting user logs into heterogeneous graphs, the "message passing" training process, and the technical pitfalls of graph data leakage. She shares how a hybrid architecture solved inference latency, delivering contextual embeddings to a downstream model.

![Image 3: Reimagining Platform Engagement with Graph Neural Networks/presentations/graph-neural-networks/en/smallimage/Mariia-Bulycheva-thumbnail-1775048997053.jpeg)](http://www.infoq.com/presentations/graph-neural-networks)

All in ai-ml-data-engFollow Topic

[Culture & Methods](http://www.infoq.com/culture-methods/ "Culture & Methods")

  • [Agile](http://www.infoq.com/agile/ "Agile")
  • [Diversity](http://www.infoq.com/diversity/ "Diversity")
  • [Leadership](http://www.infoq.com/leadership/ "Leadership")
  • [Lean/Kanban](http://www.infoq.com/lean/ "Lean/Kanban")
  • [Personal Growth](http://www.infoq.com/personal-growth/ "Personal Growth")
  • [Scrum](http://www.infoq.com/scrum/ "Scrum")
  • [Sociocracy](http://www.infoq.com/sociocracy/ "Sociocracy")
  • [Software Craftmanship](http://www.infoq.com/software_craftsmanship/ "Software Craftmanship")
  • [Team Collaboration](http://www.infoq.com/team-collaboration/ "Team Collaboration")
  • [Testing](http://www.infoq.com/testing/ "Testing")
  • [UX](http://www.infoq.com/ux/ "UX")

Featured in Culture & Methods

Celine Pypaert discusses the ubiquitous nature of open-source software and shares a blueprint for securing modern applications. She explains how to prioritize high-risk vulnerabilities using exploitability data, the role of Software Bill of Materials (SBOM), and the importance of bridging the gap between DevOps and Security through clear accountability and automated governance.

![Image 4: Empower Your Developers: How Open Source Dependencies Risk Management Can Unlock Innovation/presentations/open-source-dependencies/en/smallimage/celine-pypaert-thumbnail-1775047335370.jpeg)](http://www.infoq.com/presentations/open-source-dependencies)

All in culture-methodsFollow Topic

DevOps

  • [Infrastructure](http://www.infoq.com/infrastructure/ "Infrastructure")
  • [Continuous Delivery](http://www.infoq.com/continuous_delivery/ "Continuous Delivery")
  • [Automation](http://www.infoq.com/automation/ "Automation")
  • [Containers](http://www.infoq.com/containers/ "Containers")
  • [Cloud](http://www.infoq.com/cloud-computing/ "Cloud")
  • [Observability](http://www.infoq.com/observability/ "Observability")

Featured in DevOps

Docker Extensions boost developer speed but create a "visibility gap" by isolating telemetry. To meet enterprise needs, extensions must act as bridges to centralized platforms. This article details how to use OpenTelemetry, policy-as-code, and encryption to build secure pipelines. Learn to balance developer productivity with the governance required for scalable, compliant observability.

![Image 5: Beyond One-Click: Designing an Enterprise-Grade Observability Extension for Docker/articles/enterprise-grade-observability-extension-docker/en/smallimage/enterprise-grade-observability-extension-docker-thumbnail-1775560652994.jpg)](http://www.infoq.com/articles/enterprise-grade-observability-extension-docker)

All in devopsFollow Topic

[Events](https://events.infoq.com/ "Events")

Helpful links

  • [About InfoQ](http://www.infoq.com/about-infoq "About InfoQ")
  • [InfoQ Editors](http://www.infoq.com/infoq-editors "InfoQ Editors")
  • [Write for InfoQ](http://www.infoq.com/write-for-infoq "Write for InfoQ")
  • [About C4Media](https://c4media.com/ "About C4Media")
  • [Diversity](https://c4media.com/diversity "Diversity")

Choose your language

  • [En](http://www.infoq.com/news/2026/04/turboquant-compression-kv-cache/# "InfoQ English")
  • 中文
  • 日本
  • Fr

![Image 6: InfoQ Architect Certification - image Online InfoQ Architect Certification Join Luca Mezzalira for this 5-week online cohort. Master socio-technical architecture leadership. Register Now.](https://certification.qconferences.com/?utm_source=infoq&utm_medium=referral&utm_campaign=homepageheader_onlinecohortaprmayjun26)![Image 7: QCon AI Boston - image QCon AI Boston Learn how leading engineering teams run AI in production—reliably, securely, and at scale. Early Bird ends April 14.](https://boston.qcon.ai/?utm_source=infoq&utm_medium=referral&utm_campaign=homepageheader_qaiboston26)![Image 8: QCon San Francisco - image QCon San Francisco Learn what's next in AI and software, from teams already doing it. Early Bird ends April 14.](https://qconsf.com/?utm_source=infoq&utm_medium=referral&utm_campaign=homepageheader_qsf26)

[InfoQ Homepage](http://www.infoq.com/ "InfoQ Homepage")[News](http://www.infoq.com/news "News")Google’s TurboQuant Compression May Support Faster Inference, Same Accuracy on Less Capable Hardware

[AI, ML & Data Engineering](http://www.infoq.com/ai-ml-data-eng/ "AI, ML & Data Engineering")

Shipping Faster, Breaking More: Rethinking Delivery Systems in the Age of AI (Webinar May 28th)

Google’s TurboQuant Compression May Support Faster Inference, Same Accuracy on Less Capable Hardware

Apr 15, 2026 3 min read

by

Follow Application Consultant

#### Write for InfoQ

Feed your curiosity.Help 550k+ global

senior developers

each month stay ahead.Get in touch

Log in to listen to this article

Loading audio

Your browser does not support the audio element.

0:00 0:00

Normal 1.25x 1.5x

Like

Google Research unveiled TurboQuant, a novel quantization algorithm that compresses large language models’ Key-Value caches by up to 6x. With 3.5-bit compression, near-zero accuracy loss, and no retraining needed, it allows developers to run massive context windows on significantly more modest hardware than previously required. Early community benchmarks confirm significant efficiency gains.

While the rationale for quantization may seem logical, the difficulty is, given a number of encoding bits, to maintain the accuracy of inference-relevant computations (e.g., inner products, cosine similarity, distances) with compressed data.

The research team claims that TurboQuant can compress the KV cache down to 3.5 bits per value with near-zero accuracy loss. On standard benchmarks like LongBench and Needle in a Haystack, a 3.5-bit TurboQuant implementation matched the performance of full 16-bit precision across Gemma and Mistral models.

TurboQuant uses a two-step approach. First, data vectors are rotated (randomized Hadamard transform). This keeps key Euclidean properties (e.g., distance) while spreading out the values, removing the outlier-heavy coordinate distribution that makes low-bit quantization difficult. Post-transform, the vector coordinates follow a beta distribution that is more amenable to compression with low distortion. Second, a decade-old technique, the Quantized Johnson-Lindenstrauss (QJL) transform, is applied to remove the bias created by the first step. Post-QJL, the paper argues that inner products between quantized vectors are unbiased, computationally efficient, accurate estimators of the unquantized vectors, with the resulting effect of maintaining inference accuracy.

Early community analysis seems to confirm significant gains, albeit more modest than those reported in the paper. The Two Minute Papers analysis suggests more realistic, “real-world” improvements of 30-40% in memory reduction and processing speed:

Based on the results, we cannot conclude that every AI machine suddenly needs 6 times less RAM. No. That is a bit idealistic and only true for some corner cases. You know when you see an official benchmark of a phone battery or electric car mileage with somewhat idealized conditions? It is a bit like that.

So careful with the media hype. […] We wait for more data and analyze experiments here, to get the highest quality information.

But it’s still good. Really good! It helps most people who run AI systems with very long contexts. When you chuck in a huge PDF document, or a movie, or a huge codebase for an AI to analyze. Yes, you will be able to do that cheaper, with meaningfully less memory. Often a few gigabytes less. And I think that is absolutely amazing news.

A fundamental optimization in LLM inference is the caching of computations that are required repeatedly. This is particularly critical during autoregressive generation, where each newly generated token uses data already computed during the generation of all previous tokens. By caching these Key and Value tensors (the KV cache), the system avoids redundant, computationally expensive passes over the entire sequence history.

However, the efficiency gains offered by caching come with a significant memory cost that grows linearly with the token sequence length. For LLMs designed with long context windows, the massive VRAM footprint of the cache eventually outweighs the memory required for the model weights themselves.

For example, according to Darshan Fofadiya, AI researcher at Amazon,running a Llama 70B model with a 1M-token context window may require approximately 328GB of VRAM just for the KV cache. When compared to the 140 GB required to hold the 70B model weights in BF16, the cache becomes the primary barrier to deployment, forcing engineers into costly multi-GPU configurations. Compressed down from 16 to 3.5 bits, the cache then requires 72 GB and fits on a single H100 (80 GB HBM).

During the inference decoding phase, certain input tokens in a prompt produce KV vectors with magnitudes in the hundreds or thousands, while the majority of other tokens have values close to 0-1. In LLaMA-2-7B, for instance, the top 1% of KV cache values may have magnitudes that are 10-100x larger than the median value. This massive distribution skew makes linear 4-bit quantization impossible without specialized techniques, as the outliers stretch the quantization grid and crush the precision of normal tokens.

Generative inference with LLMs is, for relatively small batch sizes, memory-bound. With memory speeds growing slower than compute speeds, reducing the memory bottleneck (the so-called memory wall) is key to efficient inference. For short contexts, weight matrices are the dominant contributor to memory consumption. For long contexts, the KV cache becomes the main contributor. Quantization techniques for both model weights and the KV cache are thus instrumental in speeding up inference and constitute a major research topic.

About the Author

[](http://www.infoq.com/profile/Bruno-Couriol/)

#### Bruno Couriol

MSc in Telecommunications. BSc in Mathematics.

Show more Show less

#### This content is in the AI, ML & Data Engineering topic

Follow Topic

##### Related Topics:

Followers: 4084

Follow Topic

Followers: 5862

Follow Topic

Followers: 9

Follow Topic

Followers: 0

Follow Topic

Followers: 279

Follow Topic

Followers: 137

Follow Topic

* #### Popular in AI, ML & Data Engineering

* #### Related Sponsors

  • #### Related Sponsor

![Image 9: Related sponsor icon/filters:no_upscale()/sponsorship/topic/ae9df779-fe62-46d8-a42e-92795ae3c56e/promptfoo-horizontal-logo-1775562471842.png)](http://www.infoq.com/url/f/9e1e2056-ec65-4658-aaaa-50b66b2d0ee1/)Confidently test, evaluate, and red-team your LLM apps with Promptfoo — catch regressions, benchmark models, and ship high-quality AI features faster; start testing your prompts today. [Learn More](http://www.infoq.com/url/f/0ed8a8f2-ad41-400e-b24f-e10459b3993d/).

Related Content

Apr 14, 2026

Apr 13, 2026

Apr 13, 2026

Apr 07, 2026

Mar 27, 2026

Mar 26, 2026

Mar 24, 2026

Mar 23, 2026

Mar 20, 2026

Related Sponsors

The Model Context Protocol (MCP) defines a standard way for AI systems to interact with tools, data, and services. This article explains MCP’s architecture—hosts, clients, and servers—and how it enables structured, secure integrations between AI models and external systems.

System prompts define how LLM applications behave—but they are vulnerable to manipulation. This article explores prompt hardening techniques such as instruction shielding, syntax reinforcement, and layered prompting to defend AI systems against prompt injection and override attacks.

  • Sponsored by

![Image 12: Icon image/filters:no_upscale()/sponsorship/topic/ae9df779-fe62-46d8-a42e-92795ae3c56e/promptfoo-horizontal-logo-1775562471842.png)](http://www.infoq.com/url/f/9e1e2056-ec65-4658-aaaa-50b66b2d0ee1/)

Related Content

Mar 20, 2026

Mar 19, 2026

Apr 02, 2026 ![Image 13: Icon image/articles/beyond-rag-context-aware/en/smallimage/beyond-rag-context-aware-thumbnail-1774531119239.jpg)](http://www.infoq.com/articles/beyond-rag-context-aware/)

Feb 13, 2026 ![Image 14: Icon image/presentations/llm-large-scale-applications/en/smallimage/sahil-dua-thumbnail-1769590214923.jpeg)](http://www.infoq.com/presentations/llm-large-scale-applications/)

Feb 09, 2026 ![Image 15: Icon image/articles/building-llms-resource-constrained-environments/en/smallimage/building-llms-resource-constrained-environments-thumbnail-1770217548603.jpg)](http://www.infoq.com/articles/building-llms-resource-constrained-environments/)

Jan 21, 2026 ![Image 16: Icon image/articles/ai-assisted-development-series/en/smallimage/ai-assisted-development-thumb-image-1768388299998.jpg)](http://www.infoq.com/articles/ai-assisted-development-series/)

**The InfoQ** Newsletter

A round-up of last week’s content on InfoQ sent out every Tuesday. Join a community of over 250,000 senior developers. View an example

Enter your e-mail address

Select your country - [x] I consent to InfoQ.com handling my data as explained in this Privacy Notice.

We protect your privacy.

  • ##### [Claude Code Used to Find Remotely Exploitable Linux Kernel Vulnerability Hidden for 23 Years](http://www.infoq.com/news/2026/04/claude-code-linux-vulnerability/ "Claude Code Used to Find Remotely Exploitable Linux Kernel Vulnerability Hidden for 23 Years")
  • ##### [Cloudflare Introduces EmDash: TypeScript CMS Positioned as WordPress Successor](http://www.infoq.com/news/2026/04/cloudflare-emdash-wordpress/ "Cloudflare Introduces EmDash: TypeScript CMS Positioned as WordPress Successor")
  • ##### [Stateful Continuation for AI Agents: Why Transport Layers Now Matter](http://www.infoq.com/articles/ai-agent-transport-layer/ "Stateful Continuation for AI Agents: Why Transport Layers Now Matter")
  • ##### [Zendesk Says AI Makes Code Abundant, Shifting the Bottleneck to “Absorption Capacity”](http://www.infoq.com/news/2026/04/zendesk-absorption-capacity/ "Zendesk Says AI Makes Code Abundant, Shifting the Bottleneck to “Absorption Capacity”")
  • ##### [Platform Engineering: Lessons from the Rise and Fall of eBay Velocity](http://www.infoq.com/presentations/platform-engineering-lessons/ "Platform Engineering: Lessons from the Rise and Fall of eBay Velocity")
  • ##### [Lyft Scales Global Localization Using AI and Human-in-the-Loop Review](http://www.infoq.com/news/2026/04/lyft-ai-localization-pipeline/ "Lyft Scales Global Localization Using AI and Human-in-the-Loop Review")
  • ##### [Empower Your Developers: How Open Source Dependencies Risk Management Can Unlock Innovation](http://www.infoq.com/presentations/open-source-dependencies/ "Empower Your Developers: How Open Source Dependencies Risk Management Can Unlock Innovation")
  • ##### [Tiger Teams, Evals and Agents: The New AI Engineering Playbook](http://www.infoq.com/podcasts/tiger-teams-evals-agents/ "Tiger Teams, Evals and Agents: The New AI Engineering Playbook")
  • ##### [Developing Your Leadership Skills toward Principal Engineering](http://www.infoq.com/news/2026/04/leadership-skills/ "Developing Your Leadership Skills toward Principal Engineering")
  • ##### [Google’s TurboQuant Compression May Support Faster Inference, Same Accuracy on Less Capable Hardware](http://www.infoq.com/news/2026/04/turboquant-compression-kv-cache/ "Google’s TurboQuant Compression May Support Faster Inference, Same Accuracy on Less Capable Hardware")
  • ##### [Anthropic Paper Examines Behavioral Impact of Emotion-Like Mechanisms in LLMs](http://www.infoq.com/news/2026/04/anthropic-paper-llms/ "Anthropic Paper Examines Behavioral Impact of Emotion-Like Mechanisms in LLMs")
  • ##### [Reimagining Platform Engagement with Graph Neural Networks](http://www.infoq.com/presentations/graph-neural-networks/ "Reimagining Platform Engagement with Graph Neural Networks")
  • ##### [OpenTelemetry Declarative Configuration Reaches Stability Milestone](http://www.infoq.com/news/2026/04/opentelemetry-declarative-config/ "OpenTelemetry Declarative Configuration Reaches Stability Milestone")
  • ##### [New Rowhammer Attacks on NVIDIA GPUs Enable Full System Takeover](http://www.infoq.com/news/2026/04/rowhammer-attacks-nvidia/ "New Rowhammer Attacks on NVIDIA GPUs Enable Full System Takeover")
  • ##### [Beyond One-Click: Designing an Enterprise-Grade Observability Extension for Docker](http://www.infoq.com/articles/enterprise-grade-observability-extension-docker/ "Beyond One-Click: Designing an Enterprise-Grade Observability Extension for Docker")

**The InfoQ** Newsletter

A round-up of last week’s content on InfoQ sent out every Tuesday. Join a community of over 250,000 senior developers. View an example

  • Get a quick overview of content published on a variety of innovator and early adopter technologies
  • Learn what you don’t know that you don’t know
  • Stay up to date with the latest information from the topics you are interested in

Enter your e-mail address

Select your country - [x] I consent to InfoQ.com handling my data as explained in this Privacy Notice.

We protect your privacy.

**April 15 | May 7 | June 10, 2026 | Online** Architecture decisions are hard to validate while shipping. Join a **5-week online cohort** for **senior engineers, architects, and team leads** to pressure-test real decisions, apply practical frameworks, and work through challenges with a confidential peer group. Facilitated by Luca Mezzalira, Principal Architect at AWS, this cohort helps you: * Pressure-test real decisions. * Apply frameworks to real problems. * Publish on InfoQ.com and earn your certification. **RESERVE YOUR PLACE**

#### Events

April 15 | May 7 | June 10, 2026

June 1-2, 2026

November 16-20, 2026

#### Follow us on

Youtube 232K FollowersLinkedin 26K FollowersRSS 19K ReadersX 57.1k FollowersFacebook 21K LikesBluesky NewAlexa New

#### Stay in the know

The InfoQ Podcast![Image 17: The InfoQ Podcast Logo - Stay in the know](http://www.infoq.com/podcasts/)Engineering Culture Podcast![Image 18: Engineering Culture Podcast Logo - Stay in the knoww](http://www.infoq.com/podcasts/#engineering_culture)The Software Architects' Newsletter![Image 19: The Software Architects' Newsletter Logo - Stay in the know](http://www.infoq.com/software-architects-newsletter/)

General Feedback [[email protected]](mailto:[email protected]) Advertising [[email protected]](mailto:[email protected]) Editorial [[email protected]](mailto:[email protected]) Marketing [[email protected]](mailto:[email protected])

InfoQ.com and all content copyright © 2006-2026 C4Media Inc.

Privacy Notice, Terms And Conditions, Cookie Policy

Close

[BT](http://www.infoq.com/int/bt/ "bt")