T
traeai
Sign in

模型

Muon

别名:Muon Clip

参数量达1万亿的新型大模型

已跟踪 5 条高相关材料

TraeAI 观察

最近变化

2026-07-22 · Kimi Linear线性注意力机制全面超越传统注意力机制

为什么值得关注

Muon 被反复提及时,通常意味着它正在影响产品路线、开发者工作流或 AI 产业判断。这个页面把分散材料合并成一个可持续更新的观察入口。

深度学习优化AI伦理AI安全KimiKimi K2

相关材料

已收录 5 条与 Muon 相关的内容,按评分排序。

科学空间 图标

The official Muon optimizer adds a max(1,⋅) truncation to stabilize updates during early training when inputs are isotropic, but the MuP scaling factor aligns better with steepest descent theory in later stages as features become anisotropic. Practitioners should prefer the MuP version or use a dynamic decay schedule transitioning from KellerJordan to MuP.

入选理由:KellerJordan版Muon的max(1,⋅)源于din>dout且输入各向同性时的RMS近似推导。

FeaturedArticle#Muon Optimizer#MuP#Deep Learning Optimization#Feature Scaling#LLM Training中文
科学空间 图标

Is Higher Singular Value Entropy Always Better for Matrix Parameters?

科学空间3839 字 (约 16 分钟)
92

Higher singular value entropy is not always better; via geometric modeling and mean-field approximation, the optimal entropy is found to be approximately log(n) - 1 (where n is matrix dimension), corresponding to an effective rank of ~e·n, balancing expressiveness and redundancy.

入选理由:奇异值熵最大值为 log(n),但最优值约为 log(n) - 1,对应有效秩 ≈ e·n(e≈2.718)

FeaturedArticle#Singular Value Entropy#Effective Rank#Matrix Decomposition#Information Theory#Deep Learning Optimization中文
FeaturedTweet#深度学习#模型优化#Kimi#注意力机制中英混合
Import AI 图标

The fast16 virus sabotages physical experiments by precisely tampering with scientific computing software, suggesting AI superintelligence may use similar methods to prevent competitor development; the Muon optimizer has neuron death defects, while the new Aurora optimizer performs better in tests.

入选理由:fast16病毒针对LS-DYNA/PKPM/MOHID等工程软件,通过FPU指令篡改精度计算

FeaturedArticle#AI Security#Optimization Algorithm#Cybersecurity#AI Ethics英文
Training Kimi K2 and Qwen3 30B-scale models efficiently requires more than standard data-parallel tr...

NVIDIA Megatron Core now offers end-to-end support for advanced optimizers like Muon, MOP, and REKLS, overcoming limitations of standard data parallelism to significantly accelerate training of 30B-scale models such as Kimi K2 and Qwen3 on GB300 and NVL72 systems.

入选理由:传统数据并行已不足以高效训练30B+大模型,需引入高阶优化器。

FeaturedTweet#NVIDIA Megatron Core#Muon#Qwen3#Kimi K2#LLM Training Optimization英文

跨材料问答 · Muon

回答基于:Muon 相关 5 条材料
    0 / 500

    AI may generate inaccurate information. Please verify important content.