T
traeai
Sign in

概念

SFT

别名:监督微调

监督微调技术

已跟踪 9 条高相关材料

TraeAI 观察

相关材料

已收录 9 条与 SFT 相关的内容,按评分排序。

How LLMs Learn to Be Helpful (RLHF vs DPO)

How LLMs Learn to Be Helpful (RLHF vs DPO)

ByteByteGo Newsletter2425 字 (约 10 分钟)
85

本文对比RLHF与DPO两种方法,揭示大语言模型如何通过偏好学习提升帮助性,解析训练三阶段及技术局限性。

入选理由:模型训练分三阶段:预训练、监督微调(SFT)、偏好教学(RLHF/DPO)

FeaturedArticle#LLM#RLHF#DPO#模型训练英文
Fireworks AI(@FireworksAI_HQ) 图标

Fine-tuning on proprietary data is the most strategic AI advantage; prompts are easily copied, while models trained on private data are hard to replicate. OpenAI is restricting this path—companies must act now to retain SFT control.

入选理由:使用专有数据进行SFT微调可建立竞争壁垒,防止提示工程被快速复制。

FeaturedTweet#SFT#Fine-tuning#Fireworks AI#OpenAI英文
GLM 5.1 from @Zai_org is now available on @FireworksAI_HQ Training Platform across the Managed and T...

Fireworks AI 平台正式支持智谱 GLM 5.1 模型,提供 SFT/DPO 微调能力、200K 超长上下文窗口,专为长周期智能体编程微调优化,RL 训练即将上线。

入选理由:GLM 5.1 已集成至 Fireworks AI 托管与 API 训练工作流

FeaturedTweet#GLM#Fireworks AI#大模型微调#SFT#DPO中文
Personalization in the Era of LLMs - Shivam Verma, Spotify

Personalization in the Era of LLMs - Shivam Verma, Spotify

AI Engineer5271 字 (约 22 分钟)
70

Spotify builds a highly steerable personalized recommendation system for 750M users and 100M+ tracks by transforming user action sequences into vectors and then tokens, combining content representation with LLMs.

入选理由:Spotify AI基础团队通过CPT和SFT微调开源权重LLM来构建推荐系统。

FeaturedVideo#Spotify#LLM#Personalized Recommendation#User Modeling英文
its going to be a good model

its going to be a good model

eric zakariasson(@ericzakariasson)99 字 (约 1 分钟)
60

Cursor团队在v9模型训练中贡献工程进展,但补充数据效果有限。

入选理由:Cursor团队在v9 SFT和RL训练中做出重大工程贡献

FeaturedTweet#AI模型#训练数据#Cursor#SFT#RL英文
吃透大模型SFT底层机理:终结实践争议,规避无效算力

This article discusses the underlying mechanism of large model SFT (Self-Training), aiming to put an end to practice controversies and avoid fruitless computing power. By deeply understanding the SFT mechanism, engineers can more effectively utilize resources and avoid unnecessary calculations.

入选理由:SFT机制可以有效减少算力浪费,提高模型训练效率。

FeaturedArticle#large model#SFT#resource optimization中文

跨材料问答 · SFT

回答基于:SFT 相关 9 条材料
    0 / 500

    AI may generate inaccurate information. Please verify important content.