T
traeai
Sign in

概念

SFT

别名:Supervised Fine-Tuning

监督微调训练方法

已跟踪 11 条高相关材料

TraeAI 观察

相关材料

已收录 11 条与 SFT 相关的内容,按评分排序。

Fireworks AI(@FireworksAI_HQ) 图标

Fireworks AI 现已支持 DeepSeek V4 Flash 0731 模型的微调训练,提供 SFT/DPO/RL 三种训练方式,适用于编码代理和高并发场景,成本较传统方案降低 70%。

入选理由:DeepSeek V4 Flash 0731 可通过 Fireworks API 进行 SFT/DPO/RL 训练

FeaturedTweet#Fireworks AI#DeepSeek#模型微调#SFT#DPO英文
Hugging Face Journal Club: Kimi K3

Hugging Face Journal Club: Kimi K3

Hugging Face7812 字 (约 32 分钟)
85

Kimi K3通过后训练和多领域专家模型显著提升性能,接近Opus 4.8水平,采用创新的推理分层和策略蒸馏技术。

入选理由:Kimi K3后训练阶段结合SFT与RL,构建9个领域专家模型(3领域×3推理层级)

FeaturedVideo#模型训练#SFT#RL#Hugging Face#Kimi K3中英混合
How LLMs Learn to Be Helpful (RLHF vs DPO)

How LLMs Learn to Be Helpful (RLHF vs DPO)

ByteByteGo Newsletter2425 字 (约 10 分钟)
85

本文对比RLHF与DPO两种方法,揭示大语言模型如何通过偏好学习提升帮助性,解析训练三阶段及技术局限性。

入选理由:模型训练分三阶段:预训练、监督微调(SFT)、偏好教学(RLHF/DPO)

FeaturedArticle#LLM#RLHF#DPO#模型训练英文
Fireworks AI(@FireworksAI_HQ) 图标

Fine-tuning on proprietary data is the most strategic AI advantage; prompts are easily copied, while models trained on private data are hard to replicate. OpenAI is restricting this path—companies must act now to retain SFT control.

入选理由:使用专有数据进行SFT微调可建立竞争壁垒,防止提示工程被快速复制。

FeaturedTweet#SFT#Fine-tuning#Fireworks AI#OpenAI英文
GLM 5.1 from @Zai_org is now available on @FireworksAI_HQ Training Platform across the Managed and T...

Fireworks AI 平台正式支持智谱 GLM 5.1 模型,提供 SFT/DPO 微调能力、200K 超长上下文窗口,专为长周期智能体编程微调优化,RL 训练即将上线。

入选理由:GLM 5.1 已集成至 Fireworks AI 托管与 API 训练工作流

FeaturedTweet#GLM#Fireworks AI#大模型微调#SFT#DPO中文
Personalization in the Era of LLMs - Shivam Verma, Spotify

Personalization in the Era of LLMs - Shivam Verma, Spotify

AI Engineer5271 字 (约 22 分钟)
70

Spotify builds a highly steerable personalized recommendation system for 750M users and 100M+ tracks by transforming user action sequences into vectors and then tokens, combining content representation with LLMs.

入选理由:Spotify AI基础团队通过CPT和SFT微调开源权重LLM来构建推荐系统。

FeaturedVideo#Spotify#LLM#Personalized Recommendation#User Modeling英文
its going to be a good model

its going to be a good model

eric zakariasson(@ericzakariasson)99 字 (约 1 分钟)
60

Cursor团队在v9模型训练中贡献工程进展,但补充数据效果有限。

入选理由:Cursor团队在v9 SFT和RL训练中做出重大工程贡献

FeaturedTweet#AI模型#训练数据#Cursor#SFT#RL英文
吃透大模型SFT底层机理:终结实践争议,规避无效算力

This article discusses the underlying mechanism of large model SFT (Self-Training), aiming to put an end to practice controversies and avoid fruitless computing power. By deeply understanding the SFT mechanism, engineers can more effectively utilize resources and avoid unnecessary calculations.

入选理由:SFT机制可以有效减少算力浪费,提高模型训练效率。

FeaturedArticle#large model#SFT#resource optimization中文

跨材料问答 · SFT

回答基于:SFT 相关 11 条材料
    0 / 500

    AI may generate inaccurate information. Please verify important content.