building evals is hard! we're working on some skills to try to automate as much as possible. still ...

TL;DR · AI 摘要
Harrison Chase 推出 Eval Engineering Skill,通过 Harbor 工具自动化构建评估流程,但仍需人工参与。
核心要点
- 使用 Harbor 工具可自动化生成评估脚本,减少重复劳动
- 评估构建流程包含代码库分析、用户反馈迭代、结果验证三阶段
- 人工参与仍是评估优化的关键环节,自动化仅能处理基础工作
结构提纲
按章节快速跳转。
思维导图
用一张图看清主题之间的关系。
查看大纲文本(无障碍 / 无 JS 友好)
- 自动化评估构建
- 核心流程
- 代码库分析
- 用户反馈迭代
- Harbor 工具执行
- 关键工具
- Harbor 自动化平台
- 人工参与
- 结果验证
- 方向调整
金句 / Highlights
值得收藏与分享的关键句。
still requires human in the loop, but should help bootstrap
build evals (using harbor) - run evals - look at results
The skill inspects how an agent is structured, mines...
Harrison Chase on X: "building evals is hard! we're working on some skills to try to automate as much as possible. still requires human in the loop, but should help bootstrap overall flow is: - give coding agent the codebase + actual traces - iterate on eval direction with user - build evals (using harbor) - run evals - look at results, iterate with user again, repeat" / X
@hwchase17
building evals is hard! we're working on some skills to try to automate as much as possible. still requires human in the loop, but should help bootstrap overall flow is: - give coding agent the codebase + actual traces - iterate on eval direction with user - build evals (using harbor) - run evals - look at results, iterate with user again, repeat
Viv
@Vtrivedy10
Jul 22
Article
Towards Automating Eval Engineering
Today we’re releasing our Eval Engineering Skill, a skill that helps coding agents build evals using context from a repository and agent traces. The skill inspects how an agent is structured, mines...
7:28 PM · Jul 22, 2026
23K
Views
23
21
201
216