T
traeai
Sign in

概念

RL

别名:强化学习

Reinforcement Learning的缩写

已跟踪 12 条高相关材料

TraeAI 观察

相关材料

已收录 12 条与 RL 相关的内容,按评分排序。

25+ startups all solving the same missing piece

25+ startups all solving the same missing piece

Gradient Flow888 字 (约 4 分钟)
85

强化学习正成为AI基础设施核心,25家初创公司围绕模拟环境与评分系统构建工具链,解决模型可靠性难题。

入选理由:25家初创公司聚焦强化学习基础设施,解决AI模型可靠性问题

FeaturedArticle#强化学习#AI基础设施#初创公司#机器人#工业控制英文
At test time, we wrap LLMs in scaffolds that scale compute every which way -- longer chains, paralle...

斯坦福AI实验室提出Spiral方法,通过集合强化学习(set RL)和标准强化学习(RL)训练模型,使其在推理时能利用更长的链条、并行样本和聚合计算。

入选理由:Spiral方法结合集合强化学习和标准强化学习,提升模型推理能力。

FeaturedTweet#AI#强化学习#LLM#Stanford AI Lab英文
The data black hole at the center of AI

The data black hole at the center of AI

Dwarkesh Patel2576 字 (约 11 分钟)
85

AI的进展主要依赖于数据量和计算资源的增加,而非样本效率的提升,且高质量数据获取成本极高。

入选理由:AI的进步主要依赖于数据量和计算资源的增加,而非样本效率的提升。

FeaturedVideo#AI#数据#强化学习#计算资源英文
OpenAI 发布的新论文太有趣了,有点探索人性底层原理的意味。

业界研究发现在对齐大模型的时候,有个很糟糕的现象叫 emergent misalignment(涌现失调):
一个模型如果在训练时被...

OpenAI 的新论文揭示了通过强化学习对齐大模型的道德行为,发现好行为可泛化到其他领域,对抗压力下表现更稳健。

入选理由:训练模型在特定领域表现诚实、透明,可泛化到其他领域,如健康、法律等。

FeaturedTweet#OpenAI#强化学习#AI对齐#道德行为中英混合
AI HOT 精选 图标

AI中心的数据黑洞

AI HOT 精选2089 字 (约 9 分钟)
85

AI的性能提升主要依赖于数据和计算规模,而非样本效率的提升,数据需求巨大且高度专业化。

入选理由:AI的性能提升主要依赖于数据和计算规模,而非样本效率的提升。

FeaturedArticle#AI#数据#机器学习#RL#样本效率中英混合
How Cursor Ships a 1TB Model Across the World Mid-Training

How Cursor Ships a 1TB Model Across the World Mid-Training

Sequoia Capital355 字 (约 2 分钟)
85

Cursor achieves 1TB model cross-continental synchronization during training by leveraging weight change patterns in RL, reducing transmission volume by 20x and ensuring model consistency.

入选理由:RL训练中仅少量权重变化,delta压缩使传输量减少20倍。

FeaturedVideo#Model Transfer#Delta Compression#Reinforcement Learning#Distributed Training英文
#539. 手搓AlphaGo:前DeepMind科学家拆解AI围棋核心原理,以及对LLM强化学习的深远启示

Rebuilding AlphaGo: A Deep Dive into AI Go Core Principles and Implications for LLMs

跨国串门儿计划1868 字 (约 8 分钟)
85

AlphaGo uses MCTS and neural networks to achieve efficient search, showcasing the potential of reinforcement learning.

入选理由:AlphaGo 使用 MCTS 和神经网络实现高效搜索,每步都有明确监督目标。

FeaturedPodcast#AI#Reinforcement Learning#Go#Neural Networks#Search Algorithms中文
Vol.119|对话 Macaron AI 创始人 Andrew:下一代模型公司正在从 Agent 产品里长出来?

Andrew, founder of Mind Lab (Macaron AI), argues that next-generation model companies are emerging from Agent products, using LoRA reinforcement learning and continuous learning to evolve AI Agents in real-world scenarios for personalized, interactive long-term intelligence.

入选理由:Mind Lab实现了万亿参数规模的LoRA强化学习,并构建了支持DSA和MTP的LoRA RL基础设施。

FeaturedPodcast#Agent#LoRA#Reinforcement Learning#Continuous Learning#Personal AGI中文
Cursor  | The Hidden Bug in Every Large-Scale RL Run

Cursor | The Hidden Bug in Every Large-Scale RL Run

Sequoia Capital248 字 (约 1 分钟)
75

In large-scale RL training, numerical mismatches arise due to model version drift and floating-point precision differences, causing inconsistent log probabilities during inference and introducing training bias.

入选理由:在异步训练中,需重运行前向传播以生成对数概率,但相同模型版本下结果可能不同。

FeaturedVideo#Reinforcement Learning#Large Models#Numerical Stability#Training Systems#AI Systems Engineering英文
its going to be a good model

its going to be a good model

eric zakariasson(@ericzakariasson)99 字 (约 1 分钟)
60

Cursor团队在v9模型训练中贡献工程进展,但补充数据效果有限。

入选理由:Cursor团队在v9 SFT和RL训练中做出重大工程贡献

FeaturedTweet#AI模型#训练数据#Cursor#SFT#RL英文
We've gotten really really good at RL. Composer 2.5 is fighting well-above its weight class.

Very e...

We've gotten really really good at RL. Composer 2.5 is fighting well-above its weight class.

Sualeh Asif(@sualehasif996)134 字 (约 1 分钟)
50

Cursor Composer 2.5 is officially released, achieving performance breakthroughs through reinforcement learning with double free usage for one week. The new model excels at handling long-term complex tasks, and the Cursor team is collaborating with SpaceXAI to scale model sizes and compute.

入选理由:Composer 2.5采用强化学习优化,性能表现超出预期

FeaturedTweet#Cursor#Composer 2.5#Reinforcement Learning#AI Programming Tool#SpaceXAI英文

跨材料问答 · RL

回答基于:RL 相关 12 条材料
    0 / 500

    AI may generate inaccurate information. Please verify important content.