Fireworks AI(@FireworksAI_HQ)
ICYMI from a few weeks back, we compiled our learnings around how to achieve Training-Inference Pari...
7.5内容质量

TL;DR · AI 摘要
Fireworks AI 分享在 MoE 模型中实现训练与推理一致性的关键挑战:浮点加法非结合性导致数值漂移,影响模型输出一致性。
核心要点
- 浮点加法非结合性是训练推理不一致的根本原因,(a+b)+c ≠ a+(b+c)。
- 即使数学等价的核融合操作,在实际计算中仍可能因顺序不同产生数值漂移。
- 该问题在服务 Kimi K2.5 等 MoE 模型时已引发实际 parity bug,需针对性优化。
#MoE#数值稳定性#推理优化#Fireworks AI
打开原文Don’t miss what’s happening

ICYMI from a few weeks back, we compiled our learnings around how to achieve Training-Inference Parity in MoE Models. The Fundamental Issue: FP Addition Is Not Associative. (a + b) + c ≠ a + (b + c)
Quote

Fireworks AI
@FireworksAI_HQ
Apr 18
Training-Inference Parity in MoE Models: Where Numerics Drift
When Faster ≠ Identical: Numerical Pitfalls in Serving MoE Models Kernel fusions that are mathematically equivalent can still drift numerically. Here are the parity bugs we hit across both Kimi K2.5...