T
traeai
Sign in

模型

什么是 Muon

也叫:Muon Clip

参数量达1万亿的新型大模型

为什么现在值得关注?

📰 Muon 最新动态

已收录 5 篇与「Muon」相关的 AI 资讯和分析。

科学空间 图标

The official Muon optimizer adds a max(1,⋅) truncation to stabilize updates during early training when inputs are isotropic, but the MuP scaling factor aligns better with steepest descent theory in later stages as features become anisotropic. Practitioners should prefer the MuP version or use a dynamic decay schedule transitioning from KellerJordan to MuP.

入选理由:KellerJordan版Muon的max(1,⋅)源于din>dout且输入各向同性时的RMS近似推导。

FeaturedArticle#Muon Optimizer#MuP#Deep Learning Optimization#Feature Scaling#LLM Training中文
科学空间 图标

Is Higher Singular Value Entropy Always Better for Matrix Parameters?

科学空间3839 字 (约 16 分钟)
92

Higher singular value entropy is not always better; via geometric modeling and mean-field approximation, the optimal entropy is found to be approximately log(n) - 1 (where n is matrix dimension), corresponding to an effective rank of ~e·n, balancing expressiveness and redundancy.

入选理由:奇异值熵最大值为 log(n),但最优值约为 log(n) - 1,对应有效秩 ≈ e·n(e≈2.718)

FeaturedArticle#Singular Value Entropy#Effective Rank#Matrix Decomposition#Information Theory#Deep Learning Optimization中文
FeaturedTweet#深度学习#模型优化#Kimi#注意力机制中英混合
Import AI 图标

The fast16 virus sabotages physical experiments by precisely tampering with scientific computing software, suggesting AI superintelligence may use similar methods to prevent competitor development; the Muon optimizer has neuron death defects, while the new Aurora optimizer performs better in tests.

入选理由:fast16病毒针对LS-DYNA/PKPM/MOHID等工程软件,通过FPU指令篡改精度计算

FeaturedArticle#AI Security#Optimization Algorithm#Cybersecurity#AI Ethics英文
Training Kimi K2 and Qwen3 30B-scale models efficiently requires more than standard data-parallel tr...

NVIDIA Megatron Core now offers end-to-end support for advanced optimizers like Muon, MOP, and REKLS, overcoming limitations of standard data parallelism to significantly accelerate training of 30B-scale models such as Kimi K2 and Qwen3 on GB300 and NVL72 systems.

入选理由:传统数据并行已不足以高效训练30B+大模型,需引入高阶优化器。

FeaturedTweet#NVIDIA Megatron Core#Muon#Qwen3#Kimi K2#LLM Training Optimization英文

与「Muon」经常一起出现的 AI 术语。

💡 想追踪「Muon」的长期趋势?去 实体雷达 · Muon 查看详细分析和跨材料问答。

AI may generate inaccurate information. Please verify important content.