Hugging Face(@huggingface)

The same model, with the same weights, scores 62% in one agent harness and 33% in another. @adithya...

8.5内容质量
The same model, with the same weights, scores 62% in one agent harness and 33% in another.

@adithya...

TL;DR · AI 摘要

Hugging Face发布多Harness RL训练指南,使用代理提升模型性能达12%。开源代码与训练数据可复现结果。

核心要点

  • 使用代理记录token IDs和logprobs可使LFM2.5-2.6B模型性能提升12%
  • 多Harness训练使OpenCode模型准确率从34%提升至58%
  • 开源代码包含训练框架TRL、数据集和七种训练模型

结构提纲

按章节快速跳转。

  1. 揭示同一模型在不同Harness下性能差异达90%的行业痛点

  2. 通过代理层统一处理四种API格式实现跨Harness训练

  3. 基于vLLM采样记录的token级数据进行强化学习

  4. 多Harness训练使模型性能平均提升15-20%

思维导图

用一张图看清主题之间的关系。

查看大纲文本(无障碍 / 无 JS 友好)
  • 多Harness RL训练指南
    • 核心机制
      • 代理层统一API格式
    • 训练方法
      • vLLM采样记录
      • 跨Harness训练
    • 实验成果
      • 性能提升12%
      • 开源所有训练资源

金句 / Highlights

值得收藏与分享的关键句。

#强化学习#Hugging Face#多Harness训练#开源AI
打开原文

Hugging Face on X: "The same model, with the same weights, scores 62% in one agent harness and 33% in another. @adithya_s_k and the @huggingface team just released the ultimate guide to multi-harness RL, and it's one of the most practical RL write-ups this year, and everything open! The trick is … / X

Hugging Face

@huggingface

The same model, with the same weights, scores 62% in one agent harness and 33% in another.

@

adithya_s_k

and the

huggingface

team just released the ultimate guide to multi-harness RL, and it's one of the most practical RL write-ups this year, and everything open! The trick is simple. Don't touch the harness. Point it at a proxy instead of the model. The proxy speaks all four API formats coding agents use (OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, Gemini). It records the exact token ids and logprobs vLLM sampled, and you train on that. You don't change a single line of Claude Code, Codex or OpenCode. Results: 🔹 Trained across 4 harnesses at once, LFM2.5-2.6B by

liquidai

went from 42% to 54% 🔹 31% fewer tool calls, thanks to a small bonus for solving tasks in fewer steps 🔹 Training in OpenCode alone took OpenCode from 34% to 58%, but the multi-harness model improved everywhere They also tried the shortcut everyone reaches for: fine-tune on 3,189 successful rollouts from Qwen3.8-27B. Imitation plateaued at 47.5%, below both RL runs. Copying a bigger model doesn't get you there. Practice does. The best part is that everything is open: the capture proxy in OpenEnv, the trainer in TRL, the tasks, the SFT data, the training code and all seven trained models. Agents will run in dozens of harnesses. Now open models can be trained for each of them, by anyone. Read it here 👇

huggingface.co/spaces/FineEnv…

2:51 PM · Oct 2, 2026

·

85.2K

Views

127

131

954

867