The Algorithmic Bridge

如何一个小公司成功地对抗AI垃圾内容

8.5内容质量
如何一个小公司成功地对抗AI垃圾内容

TL;DR · AI 摘要

这篇文章讨论了一个小型创业公司如何有效地检测AI垃圾内容。

核心要点

  • Pangram Labs通过最大化确保捕获内容为AI来提高检测准确性。
  • Pangram Labs的技术基于Voltaire的格言:完美是敌人。
  • Pangram Labs的技术提高了检测AI垃圾内容的精度。

结构提纲

按章节快速跳转。

  1. 文章介绍了一个小型创业公司如何有效地检测AI垃圾内容。

  2. AI垃圾内容难以检测但容易识别。

  3. 现有的AI检测工具倾向于误判人类写作。

  4. Pangram Labs采用Voltaire的格言:完美是敌人,通过最大化确保捕获内容为AI来提高检测准确性。

  5. Pangram Labs的技术提高了检测AI垃圾内容的精度。

  6. 通过图表展示了Pangram Labs的技术效果。

思维导图

用一张图看清主题之间的关系。

查看大纲文本(无障碍 / 无 JS 友好)
  • AI检测
    • Pangram Labs
      • 技术
      • 效果

金句 / Highlights

值得收藏与分享的关键句。

  • Pangram Labs的技术基于Voltaire的格言:完美是敌人,通过最大化确保捕获内容为AI来提高检测准确性。
    ⬇︎ 下载 PNG𝕏 分享到 X
#AI检测#Pangram Labs
打开原文
Image 1
Image 1

嘿,Alberto👋 每周我都会发布一篇长篇的AI分析文章,涵盖文化、哲学和商业领域,供《算法桥》杂志阅读。付费订阅者还会收到每周一的实用指南和每周五的新闻评论。我也会偶尔发表一些额外的文章。如果你希望成为付费订阅者,请点击这里:

全盘披露:这不是一个付费赞助。

Image 2
Image 2

你注意到的第一件事是AI垃圾内容读起来就像是一般的在线写作。

从这个意义上说,它属于一种长期存在的污染,包括假币、掺假食品、宣传以及以教育内容名义出现的短视频。

然而,与其它污染不同的是,AI垃圾内容难以区分但并不难检测。

重要的是要暂停一下,理解“区分”意味着将某物从其周围环境隔离出来,而“检测”则意味着知道某物的存在。检测器并不关心除了目标事物之外可能还存在什么额外的东西。

AI垃圾内容很容易检测并因此可以根除。任何标准分类器都可以做到这一点,我也能做到。只要你愿意尝试,你也可以做到。实际上,AI垃圾内容之所以能够在互联网上广泛传播,是因为检测器太擅长捕捉机器了。他们甚至捕捉到了人类。问题在于:人类写作就是那些检测器在大规模渔网中捕获的额外东西,尽管大家都清楚它们在那里。

直到现在。登场的是Pangram实验室。

理解Pangram处理AI垃圾内容的方法——关键在于它与其他大多数AI检测工具的区别——可以通过伏尔泰的名言来体现:完美是恶的敌人。

与其试图通过检测完美地识别出什么是垃圾内容和什么是非垃圾内容来发起一场全面战争,他们明智地选择了最大化他们在最重要的战斗中的胜算:确保他们捕捉到的一切都是真正的AI。或者,用技术术语来说,他们想要实现近乎零的“假阳性率”(FPR)。你可以通过查看这些图表来理解为什么这很重要:

Image 3
Image 3

“ICLR提交的AI生成摘要所占比例随年份变化的图表。”来源:Pangram

Image 4
Image 4

“联邦民事投诉中被归类为包含AI生成文本的比例。”来源:Shah & Levy, 2026

有了Pangram,你可以几乎100%确定“后AI”内容中AI生成的比例,因为Pangram不会在“前AI”内容上失败。

这就是让Pangram能够做出像这样的声明的原因:“ICLR 2026会议中有21%,即15,899份评审意见是完全由AI生成的,并且超过一半的评审意见都涉及某种形式的AI参与,无论是AI编辑、辅助还是完全由AI生成。” 或者,关于亚马逊产品评论:“我们研究的总前端产品评论中,有3%是高度自信地认为是由AI生成的。” 或者,2024年8月,“每天有60,000篇AI生成的新闻文章”被发布(想象一下现在的数字)。

Image 5
Image 5

_“EditLens 预测结果的分布于 ICRA 2026 评论。” 来源:Pangram_

Image 6
Image 6

_“按发布者组织的 AI 内容发布量” + “根据国家生成的 AI 文章图。” 来源:Pangram_

我应当指出的是,Pangram 的假正类率(FPR)非常低,但仍然不是零。他们声称其 FPR 在测试集文档上为每 10,000 个文档中有一个,而在未使用的来自 ArXiv 的科学论文上为每 100,000 个文档中有一个。

自那些嘲讽 AI 探测器错误地将美国宪法或《圣经》中的某段落标记为 AI 写作的文章以来,Pangram 已经取得了长足的进步。那种荒谬性正是 Pangram 能够避免的。他们坚持遵循威廉·黑斯廷斯于 1765 年提出的原则,该原则适用于所有现代法律体系:让 10 名罪犯逃脱比让一名无辜者受罚更好。

我在之前曾以完全相同的方式捍卫过这一理念。我认为这是更好的做法:尽一切可能避免惩罚那些试图在给予他们理由变得腐败的世界中公平行事的人。然而,直到现在,我一直目睹着复杂的探测器在对抗 AI 流氓时被击败了——试图赢得每一场比赛。很难同时做到既不会误判人类创作的内容为 AI 写作(假阳性),也不会误判由 AI 创作的内容为人类创作(假阴性)。

如果你试图做到这两点,你将什么也做不了。你必须妥协。在我看来,Pangram 的决定性妥协就是人类对网络重新夺回战争的第一次成功的进攻。

但是等等,Pangram 还声称其近似零的假阴性率(FNR),即 AI 写作逃过检测的比例。芝加哥大学独立研究人员 已证实了这一点。似乎 Pangram 结束了这场战争!这看起来是什么样的交易?让我们试着理解一下发生了什么。

这意味着,从表面上看,Pangram 正像其他所有 AI 探测公司一样,正在追求整个市场。这似乎破坏了我刚刚讲述的关于赢得战斗而不失去战争的故事,并强调了妥协的重要性。夸大黑斯廷斯的交易可能会让我成为我批评其他 AI 探测公司营销策略中所批评的那种懒惰的一部分。

问题在于,探测器无法实现接近零的真阴性率(TNR)。而 FPR 总是真实的——你可以通过回到过去,那时 AI 还不存在,来确信一个探测工具的有效性。然而,对于 FNR,你需要做一些会削弱你标记所有 AI 内容的能力的技巧。

当 Pangram 和验证它们的研究人员测量假阴性时,他们只能这样做:他们使用 AI 模型生成文本,也许使用有限的“AI 人性化”功能,然后通过探测器运行这些文本,并计数漏报的数量。与这个基准相比,Pangram 的假阴性率确实非常接近于零,这可以用于公关目标和论文撰写。

问题在于,这样的测量告诉我们关于真实世界的情况几乎 _毫无意义_ :实验室控制实验中 AI 存在的分布不一定与真实世界中 AI 存在的分布有任何相似之处。相反,他们接受无法在野外准确测量 FNR 的事实,从而采用一种较弱的 FNR 定义并因此做出一个较弱的声明。这没问题——AI 探测器无法做得更好——但这与说我们从未错过标记 AI 或人类的内容有着非常不同的含义。当你不知道 AI 在现实中是如何使用的!

例如——允许我提供一个轶事数据点——我经常愚弄 Pangram。

我清楚地愚弄了每一个探测器。我很容易做到,无需特殊工具、人类化软件或精心设计的对抗性技巧。我只是给文本赋予我的风格,添加一些词语或重写几句话,让 AI 写剩下的部分。这更多是在“改变 X 个词”,而不是拥有一个能够嗅出这些工具如何欺骗的“嗅觉”。有时我会写 10%,而 AI 写 90%,结果还不错。有时情况正好相反。我这样做是为了在实验写作中测试 AI 的极限。但我也会在私下这样做,以证明测量的 FNR 是无用的。每次,Pangram 都会承认:100% 人类写作。

我的论点并不是说我 _知道_ 如何愚弄 AI 探测器,而是任何有竞争力的作家都会。事实上,我怀疑大多数人会尝试;我认为这是实践中发生的事情的正确方式。很少有人如此懒惰——除了可能 LinkedIn 的意见领袖——以至于他们会自动化整个管道并检查显而易见的结构和语法提示(例如,“这不是 X,而是 Y,三元组,喉咙清嗓过渡等。”)。

That’s precisely what FNR benchmarks consistently miss. They test pure AI text against pure human text; two categories that, increasingly, nobody operates in. There are as many ways to blend AI into your writing process as there are writers; the real world is a gradient.

You don’t get to claim the hard victory when your tool can’t measure the hard cases. To claim near-zero _true_ false negatives in these conditions is like claiming you’ve caught every target fish in a lake with your large-scale drift net when your evidence is that you’ve caught every target fish _you put there yourself._

This limitation is what makes Pangram more valuable rather than less so.

I will forgive their PR team for missing the mark on FNRs because in practice, Pangram chooses to optimize for the battle where victory is verifiable. They know what aspect of AI detection is doing the heavy lifting and act accordingly. Pangram will actually lean on “human-written” when unsure. That’s the right engineering decision and, I’d argue, the right ethical one too. That’s also the real win: not so much a zero score on both FPR and FNR, as having achieved the lowest FPR.

(This section was ~50% written by Claude with a few edits on my end; Pangram’s verdict on this article, which I share at the end, will tell you how much weight to put on those FNR numbers.)

I’ve criticized AI detection tools before. I’ve also criticized a naive “AI;DR” approach—it’s AI, so I didn’t read it—by which people assume they’re good enough at distinguishing AI so that they can decide with a glimpse. Those who’ve read me for a while know that I’m skeptical about our ability to do anything of lasting effect to prevent a total AI-induced destruction of the digital commons.

Well, I might have been right in the specifics but I was wrong in my conclusion. There is hope.

The reason Pangram is different is that they allow us, with a really high confidence, to sort out AI slop. If you trust the tool—I do—you can be sure that, when it flags something as AI-generated, it is AI-generated. This removes the main reason why the other detectors were worse than nothing: the liar’s dividend. No one can hide under the pretext that “AI detectors often flag as AI what’s human-written.” To this end, they’ve just announced the next step: a Chrome extension that works across platforms (X, Substack, Medium, LinkedIn, Reddit, etc.).

This is, essentially, humanity delaying a definitive but dubious victory against AI slop in exchange for winning with certainty the most urgent battle; a tradeoff I’ll take every time. By virtue of being good, Pangram is not perfect (which is the other side of Voltaire’s maxim) and that’s ok. Is it so terrible that some guilty individuals will escape our judgment? No, when you consider that not pursuing them ensures that almost no innocent people are harmed.

Besides, you can always trust _your instincts_. You may have heuristics, rules of thumb that work in distinguishing cues and signs that Pangram has not encoded. That’s ok, use them; we are not perfect detectors of true FNR but, with enough training, definitely better than Pangram is. You should not blindly trust that something is human-made just because Pangram says so.

That said, I don’t support a witch hunt. I understand the hate and the frustration because we’re all drowning in AI slop, but public shame and cancellation are something that’s better left behind. This article is less a condemnation of AI writing itself and the people who engage with it, and more a defense of individual power to decide what you want your info diet to be made of.

The stance I support is a sort of “functional” AI;DR. Whereas AI;DR is a blanket rejection, a functional version combines the use of the best AI detectors and your natural skills. In fact, not only do I support this, but I encourage you to do it: the only way to clean up the entire digital town is for each of us to clean the sidewalk in front of our own digital homes. Take local action to see global results.

_Hey, Pangram, how did I do?_

Image 7
Image 7

_Pangram’s score for this article. Source: Pangram_