T
traeai
Sign in

概念

PPO

别名:Proximal Policy Optimization

近端策略优化算法,用于在线对齐训练。

已跟踪 2 条高相关材料

TraeAI 观察

相关材料

已收录 2 条与 PPO 相关的内容,按评分排序。

机器人运控训练步入分钟级时代!清华AIR开源UniLab:3分钟训好人形,速度暴涨10倍,Mac上也能跑

Tsinghua University's AIR DISCOVER Lab open-sources UniLab, achieving 3-10x end-to-end training speedup through heterogeneous architecture, supporting local training on Mac and enabling humanoid robot training in minutes, marking the arrival of the minute-level era for embodied intelligence.

入选理由:UniLab采用CPU仿真+GPU训练的异构架构,实现3-10倍端到端训练加速。

FeaturedArticle#Robotics#Reinforcement Learning#Embodied Intelligence#Open Source#Heterogeneous Computing中文
Apple Machine Learning Research 图标

Understanding Alignment in Multimodal LLMs: A Comprehensive Study

Apple Machine Learning Research519 字 (约 3 分钟)
85

苹果提出多模态大模型对齐新方法BDHS,无需额外标注即可提升模型性能,揭示离线与在线对齐方法结合的有效性。

入选理由:BDHS方法通过偏差驱动幻觉采样,无需外部模型或人工标注

FeaturedArticle#多模态LLM#对齐方法#Apple#DPO#PPO英文

跨材料问答 · PPO

回答基于:PPO 相关 2 条材料
    0 / 500

    AI may generate inaccurate information. Please verify important content.