cohere(@cohere)
For real agentic workloads (North), short-context calibration wasn't enough. We calibrated AWQ on lo...
8.5内容质量

TL;DR · AI 摘要
Cohere分享了针对长上下文工作负载优化AWQ校准的技术实践。
核心要点
- 短上下文校准不足以满足复杂工作负载需求。
- 通过token masking排除重复模板提升校准效果。
- 引入量化感知蒸馏(QAD)匹配BF16模型质量。
#机器学习#优化#Cohere#AWQ
打开原文Cohere on X: "For real agentic workloads (North), short-context calibration wasn't enough. We calibrated AWQ on long internal agentic traces (up to 64k tokens) and added token masking in llm-compressor to exclude repetitive chat templates/tool descriptions from calibration stats. Plus QAD https://t.co/n8riV16WKc" / X
Don’t miss what’s happening

For real agentic workloads (North), short-context calibration wasn't enough. We calibrated AWQ on long internal agentic traces (up to 64k tokens) and added token masking in llm-compressor to exclude repetitive chat templates/tool descriptions from calibration stats. Plus QAD (quant-aware distillation) to close the last gap — matching the quality of our BF16 MoE model with W4A8.
·
1
5