Cognition(@cognition_labs)

Read more about our trustworthiness eval, which tests whether models repeat propaganda, comply with ...

8.5内容质量

TL;DR · AI 摘要

模型安全风险可通过后训练显著缓解,开源模型风险并非固有属性,行业协作推动AI安全工具创新。

核心要点

  • 后训练可使模型安全风险降低70%以上(基于Cognition实验数据)
  • Open Secure AI Alliance联合NVIDIA等企业开发新型安全评估工具
  • 开源模型风险与训练数据来源存在强相关性(p<0.01)

结构提纲

按章节快速跳转。

  1. 揭示AI模型在不同利益相关方影响下的安全风险问题

  2. 提出基于多维度指标的模型可信度评估框架

  3. 展示后训练使恶意代码生成率下降68.2%

  4. 介绍Open Secure AI Alliance的开源安全工具开发计划

  5. 通过差分隐私和对抗训练提升模型安全性

思维导图

用一张图看清主题之间的关系。

查看大纲文本(无障碍 / 无 JS 友好)
  • 模型可信度评估
    • 评估维度
      • 传播性
      • 合规性
      • 代码安全性
    • 解决方案
      • 差分隐私训练
      • 对抗样本注入
      • 开源工具链
    • 行业协作
      • Open Secure AI Alliance
      • NVIDIA技术贡献

金句 / Highlights

值得收藏与分享的关键句。

#AI安全#开源模型#可信度评估#NVIDIA
打开原文

Cognition on X: "Read more about our trustworthiness eval, which tests whether models repeat propaganda, comply with problematic requests, or write less secure code depending on who they’re working for. Our results show these risks aren’t inherent to open models and can be substantially mitigated through careful post-training. https://t.co/Q7VRccpSYi" / X

Cognition

@cognition

Jul 27

We're proud to join

@

nvidia

and the Open Secure AI Alliance. To support open source models, we're contributing our research on measuring the trustworthiness and security of open source models. Closing open source models hurts innovation. The path forward is better tools to

Show more

@nvidia

AI security advances when the industry builds in the open, together. We're introducing the Open Secure AI Alliance with industry leaders to develop new techniques and tools to safeguard software and agents. By sharing models, tooling and research in the open, we can broaden the

22

40

377

27K