We compressed Hy4-preview from 1.5TB to ~200GiB GGUF and it still works well ! Meet MIX-STQ1_0.The...

TL;DR · AI 摘要
腾讯团队通过MIX-STQ1_0方法将Hy4-preview模型压缩至200GiB,保持高性能且误差变化小。
核心要点
- 使用MIX-STQ1_0量化方法可将1.5TB模型压缩至200GiB
- 最低位宽达1.31-bit STQ1_0,最高2.06-bit IQ2_XXS
- MCP Atlas基准测试误差仅下降0.5个百分点
结构提纲
按章节快速跳转。
- §压缩成果
展示Hy4-preview从1.5TB压缩至200GiB GGUF的成果
通过校准数据动态选择各层位宽的量化策略
- ›性能验证
在多个基准测试中保持与BF16相近的准确率
- ·技术细节
位宽范围从1.31-bit到2.06-bit的混合量化方案
思维导图
用一张图看清主题之间的关系。
查看大纲文本(无障碍 / 无 JS 友好)
- 模型压缩技术
- MIX-STQ1_0方法
- 动态位宽分配
- 校准数据驱动
- 1.31-2.06 bit范围
- 应用案例
- Hy4-preview压缩
- 200GiB GGUF格式
- 多基准测试验证
金句 / Highlights
值得收藏与分享的关键句。
压缩后模型大小降低至原尺寸的13.3%,但MCP Atlas基准测试误差仅下降0.5个百分点
采用动态位宽分配策略,部分层低至1.31-bit STQ1_0,部分层保持2.06-bit IQ2_XXS
Hy4-preview参数量达770B,支持1M上下文长度的超大规模模型
Tencent Hy on X: "We compressed Hy4-preview from 1.5TB to ~200GiB GGUF and it still works well ! Meet MIX-STQ1_0.The trick isn’t just going low, it’s deciding where: calibration data picks each layer’s bit-width, some down to 1.31-bit STQ1_0, some up to 2.06-bit IQ2_XXS. Same budget, lower error.… / X
Tencent Hy
@TencentHunyuan
We compressed Hy4-preview from 1.5TB to ~200GiB GGUF and it still works well ! Meet MIX-STQ1_0.The trick isn’t just going low, it’s deciding where: calibration data picks each layer’s bit-width, some down to 1.31-bit STQ1_0, some up to 2.06-bit IQ2_XXS. Same budget, lower error. Accuracy barely moves vs BF16 📊 MCP Atlas 83.7→83.2 📊 SWE-Bench multi 82.9→81.3 📊 MRCR 81.3→81.1 📊 IFBench 73.5→72.5 See the details on HF : AngelSlim/Hy4-preview-GGUF Weights & low-bit GGUFs 👇
huggingface.co/AngelSlim/Hy4-…
#LLM
#Quantization
#llamacpp
#Hy
Aug 28
🚀 Hy4 preview is here. 770B, 49B active, 1M context. Built for productivity. Open source frontier. Consistent affordable price. Use it. Tell us what breaks. More on Hy blog:
hy.tencent.ai/research/hy4-p…
HuggingFace:
huggingface.co/tencent/Hy4-pr…
Github:
github.com/Tencent-Hunyua…
5:31 AM · Aug 29, 2026
341.3K
Views
78
178
1.7K
553