T
traeai
Sign in

概念

LLM-as-a-Judge

别名:LLM评估系统

基于大模型的自动化评估框架

已跟踪 7 条高相关材料

TraeAI 观察

相关材料

已收录 7 条与 LLM-as-a-Judge 相关的内容,按评分排序。

How to Evaluate AI Agents with an LLM-as-a-Judge Harness in Python

How to Evaluate AI Agents with an LLM-as-a-Judge Harness in Python

freeCodeCamp.org2416 字 (约 10 分钟)
85

本文提供本地化AI代理评估框架,结合LLM作为裁判与规则检查,使用LangChain、Ollama等工具实现零API成本测试。

入选理由:使用LLM-as-a-judge与规则检查双重机制评估AI代理输出

FeaturedArticle#AI评估#LLM#Python#LangChain#Ollama英文
Databricks 图标

AI-Enabled Advisory Services for Higher Education

Databricks1365 字 (约 6 分钟)
85

Databricks通过GenAI和LLM技术解决高校咨询电话质量评估难题,实现规模化转录、评分与洞察提取。

入选理由:使用OpenAI Whisper等基础模型可提升电话转录准确率60%以上

FeaturedArticle#AI#高等教育#Databricks#GenAI#LLM英文
Machine Learning Mastery 图标

LLM Evaluation Frameworks Compared: How to Actually Measure What Your Model Does

Machine Learning Mastery4572 字 (约 19 分钟)
85

LLM评估框架存在可测量偏差,RAGAS/DeepEval/Promptfoo各有适用场景,需结合生产监控工具实现完整评估体系。

入选理由:RAGAS/DeepEval/Promptfoo三框架分别适用于不同评估场景,成熟团队常并行使用

FeaturedArticle#LLM#评估框架#RAGAS#DeepEval#Promptfoo英文
Presentation: Powering the Future: Building Your GenAI Infrastructure Stack

Intuit scaled GenAI development across 8,000+ developers with 3,500+ production experiments using the GenOS platform and 'fixed, flexible, free' framework, featuring LLM-as-a-judge evaluation and Agent-friendly API design.

入选理由:Intuit采用"fixed, flexible, free"三层框架设计GenOS平台,fixed层提供标准化基础设施,flexible层支持业务定制,free层鼓励创新实验

FeaturedArticle#AI Agent#GenAI Infrastructure#Intuit#LLM Evaluation#Platform Engineering英文

跨材料问答 · LLM-as-a-Judge

回答基于:LLM-as-a-Judge 相关 7 条材料
    0 / 500

    AI may generate inaccurate information. Please verify important content.