cohere(@cohere)

For real agentic workloads (North), short-context calibration wasn't enough. We calibrated AWQ on lo...

8.5内容质量
For real agentic workloads (North), short-context calibration wasn't enough. We calibrated AWQ on lo...

TL;DR · AI 摘要

Cohere分享了针对长上下文工作负载优化AWQ校准的技术实践。

核心要点

  • 短上下文校准不足以满足复杂工作负载需求。
  • 通过token masking排除重复模板提升校准效果。
  • 引入量化感知蒸馏(QAD)匹配BF16模型质量。
#机器学习#优化#Cohere#AWQ
打开原文

Cohere on X: "For real agentic workloads (North), short-context calibration wasn't enough. We calibrated AWQ on long internal agentic traces (up to 64k tokens) and added token masking in llm-compressor to exclude repetitive chat templates/tool descriptions from calibration stats. Plus QAD https://t.co/n8riV16WKc" / X

Don’t miss what’s happening

Image 1: Square profile picture
Image 1: Square profile picture

Cohere

@cohere

For real agentic workloads (North), short-context calibration wasn't enough. We calibrated AWQ on long internal agentic traces (up to 64k tokens) and added token masking in llm-compressor to exclude repetitive chat templates/tool descriptions from calibration stats. Plus QAD (quant-aware distillation) to close the last gap — matching the quality of our BF16 MoE model with W4A8.

Image 2: Image
Image 2: Image

8:38 PM · Apr 22, 2026

·

793 Views

1

5