Skip to content
Guides9 min read

Оценка ИИ-агентов: человеческая разметка трасс агентов

Как выстроить человеческую оценку вывода ИИ-агента в Potato: от импорта трасс через устройство схем на уровне трассы и шагов до разбора выгруженных результатов.

Potato Team

Автоматические метрики говорят, как часто агент справляется, но не почему он провалился. Человеческая оценка показывает шаг, на котором рассуждение свернуло не туда. В этом руководстве мы собираем полный конвейер человеческой оценки ИИ-агентов на возможностях разметки агентов в Potato.

The agent evaluation pipeline: raw traces are converted, displayed as cards, annotated per step and per trace, then exported as structured dataFrom raw traces to structured judgments

Мы будем оценивать агента в стиле ReAct, который отвечает на вопросы поиском в вебе. К концу у вас будет:

  • Импортированы трассы агента в Potato
  • Настроено отображение трассы агента
  • Заданы схемы разметки и по трассе целиком, и по шагам
  • Проведена задача разметки
  • Выгружены структурные результаты оценки

Отображение трасс агента в Potato рисует каждый шаг карточкой с цветовой маркировкой для мыслей, действий и наблюдений:

Agent trace display in Potato

Что понадобится

Установите Potato 2.3.0 или новее:

bash
pip install --upgrade potato-annotation

Понадобятся и сами трассы агента для оценки. В руководстве используются трассы ReAct, но тот же подход работает для любого из 15 поддерживаемых форматов.


Шаг 1: подготовьте трассы агента

Трасса ReAct — это последовательность шагов «мысль», «действие» и «наблюдение». Вот минимальный пример. Создайте файл data/traces.jsonl:

json
{
  "trace_id": "trace_001",
  "task": "What is the capital of Australia and what is its population?",
  "trace": [
    {"type": "thought", "content": "I need to find the capital of Australia and its population. Let me search for this."},
    {"type": "action", "content": "search", "action_input": "capital of Australia"},
    {"type": "observation", "content": "Canberra is the capital city of Australia. It was chosen as the capital in 1908 as a compromise between Sydney and Melbourne."},
    {"type": "thought", "content": "The capital is Canberra. Now I need to find its population."},
    {"type": "action", "content": "search", "action_input": "Canberra population 2025"},
    {"type": "observation", "content": "Canberra has a population of approximately 467,000 as of 2025, making it the eighth-largest city in Australia."},
    {"type": "thought", "content": "I now have both pieces of information. The capital of Australia is Canberra with a population of approximately 467,000."},
    {"type": "action", "content": "finish", "action_input": "The capital of Australia is Canberra, with a population of approximately 467,000 as of 2025."}
  ],
  "ground_truth": "Canberra, approximately 467,000"
}

Каждая строка файла JSONL — одна полная трасса агента. Поле trace содержит пошаговый лог. Поле task — то, что агента просили сделать.

Замечания о форматах трасс

У трасс с вызовом функций OpenAI формат выглядит иначе:

json
{
  "trace_id": "oai_001",
  "task": "Find cheap flights from NYC to London",
  "messages": [
    {"role": "user", "content": "Find cheap flights from NYC to London"},
    {"role": "assistant", "content": null, "tool_calls": [{"function": {"name": "search_flights", "arguments": "{\"from\": \"NYC\", \"to\": \"LHR\"}"}}]},
    {"role": "tool", "name": "search_flights", "content": "{\"flights\": [{\"airline\": \"BA\", \"price\": 450}, {\"airline\": \"AA\", \"price\": 520}]}"},
    {"role": "assistant", "content": "I found flights from NYC to London. The cheapest is British Airways at $450."}
  ]
}

Эти различия берёт на себя конвертер Potato. Вам нужно лишь указать правильное имя конвертера.


Шаг 2: создайте конфигурацию проекта

Создайте config.yaml:

yaml
annotation_task_name: "ReAct Agent Evaluation"
task_dir: "."
 
data_files:
  - "data/traces.jsonl"
 
item_properties:
  id_key: trace_id
  text_key: task
 
# --- Agentic annotation settings ---
agentic:
  enabled: true
  trace_converter: react
  display_type: agent_trace
 
  agent_trace_display:
    colors:
      thought: "#6E56CF"
      action: "#3b82f6"
      observation: "#22c55e"
      error: "#ef4444"
    collapse_observations: true
    collapse_threshold: 400
    show_step_numbers: true
    show_timestamps: false
    render_json: true
    syntax_highlight: true

Это говорит Potato:

  1. Загрузить трассы из data/traces.jsonl
  2. Разобрать поле trace конвертером ReAct
  3. Показывать трассы через отображение трассы агента с цветными карточками шагов

Шаг 3: продумайте схемы разметки

Оценке агента обычно нужны и суждения на уровне трассы (справился ли агент?), и на уровне шага (верен ли каждый шаг?). Добавим и то и другое.

Добавьте в config.yaml следующее:

yaml
annotation_schemes:
  # --- Trace-level schemas ---
 
  # 1. Task success (the most important metric)
  - annotation_type: radio
    name: task_success
    description: "Did the agent successfully complete the task?"
    labels:
      - "Success"
      - "Partial Success"
      - "Failure"
    label_requirement:
      required: true
    sequential_key_binding: true
 
  # 2. Answer correctness (if the task has a ground truth)
  - annotation_type: radio
    name: answer_correctness
    description: "Is the agent's final answer factually correct?"
    labels:
      - "Correct"
      - "Partially Correct"
      - "Incorrect"
      - "Cannot Determine"
    label_requirement:
      required: true
 
  # 3. Efficiency rating
  - annotation_type: likert
    name: efficiency
    description: "Did the agent use an efficient path to the answer?"
    labels:
      1: "Very Inefficient (many unnecessary steps)"
      3: "Average"
      5: "Optimal (no wasted steps)"
 
  # 4. Free-text notes
  - annotation_type: text
    name: evaluator_notes
    description: "Any additional observations"
    label_requirement:
      required: false
 
  # --- Step-level schemas ---
 
  # 5. Per-step correctness
  - annotation_type: trajectory_eval
    name: step_correctness
    steps_key: agentic_steps
    description: "Was this step correct and useful?"
    correctness_options:
      - "Correct"
      - "Partially Correct"
      - "Incorrect"
      - "Unnecessary"
 
  # 6. Per-step error type (only shown when step is not correct)
  - annotation_type: trajectory_eval
    name: error_type
    steps_key: agentic_steps
    description: "What type of error occurred?"
    correctness_options:
      - "Wrong tool/action"
      - "Wrong arguments"
      - "Hallucinated information"
      - "Reasoning error"
      - "Redundant step"
      - "Premature termination"
      - "Other"

Это покрывает оба конца разбора. Метка успеха или провала и оценка правильности ответа дают верхнеуровневые числа. Оценка эффективности позволяет сравнивать стратегии. А пошаговые оценки с таксономией ошибок, которая появляется только при пометке шага неверным, показывают, где всё сломалось.


Шаг 4: настройте вывод и запустите сервер

Добавьте в config.yaml настройки вывода:

yaml
output_annotation_dir: "output/"
export_annotation_format: "jsonl"
 
# Optional: also export to Parquet for analysis
parquet_export:
  enabled: true
  output_dir: "output/parquet/"
  compression: zstd

Полный config.yaml для справки:

yaml
annotation_task_name: "ReAct Agent Evaluation"
task_dir: "."
 
data_files:
  - "data/traces.jsonl"
 
item_properties:
  id_key: trace_id
  text_key: task
 
agentic:
  enabled: true
  trace_converter: react
  display_type: agent_trace
  agent_trace_display:
    colors:
      thought: "#6E56CF"
      action: "#3b82f6"
      observation: "#22c55e"
      error: "#ef4444"
    collapse_observations: true
    collapse_threshold: 400
    show_step_numbers: true
    render_json: true
    syntax_highlight: true
 
annotation_schemes:
  - annotation_type: radio
    name: task_success
    description: "Did the agent successfully complete the task?"
    labels: ["Success", "Partial Success", "Failure"]
    label_requirement:
      required: true
    sequential_key_binding: true
 
  - annotation_type: radio
    name: answer_correctness
    description: "Is the agent's final answer factually correct?"
    labels: ["Correct", "Partially Correct", "Incorrect", "Cannot Determine"]
    label_requirement:
      required: true
 
  - annotation_type: likert
    name: efficiency
    description: "Did the agent use an efficient path?"
    labels:
      1: "Very Inefficient"
      3: "Average"
      5: "Optimal"
 
  - annotation_type: text
    name: evaluator_notes
    description: "Any additional observations"
    label_requirement:
      required: false
 
  - annotation_type: trajectory_eval
    name: step_correctness
    steps_key: agentic_steps
    description: "Was this step correct?"
    correctness_options: ["Correct", "Partially Correct", "Incorrect", "Unnecessary"]
 
  - annotation_type: trajectory_eval
    name: error_type
    steps_key: agentic_steps
    description: "Error type"
    correctness_options:
      - "Wrong tool/action"
      - "Wrong arguments"
      - "Hallucinated information"
      - "Reasoning error"
      - "Redundant step"
      - "Premature termination"
      - "Other"
output_annotation_dir: "output/"
export_annotation_format: "jsonl"
 
parquet_export:
  enabled: true
  output_dir: "output/parquet/"
  compression: zstd

Запустите сервер:

bash
potato start config.yaml -p 8000

Откройте в браузере http://localhost:8000.


Шаг 5: как идёт разметка

Открыв трассу, разметчик видит:

  1. Описание задачи сверху (исходный запрос пользователя)
  2. Карточки шагов со всей трассой агента, с цветовой маркировкой по типу:
    • Фиолетовые карточки — мысли и рассуждения
    • Синие — действия и вызовы инструментов
    • Зелёные — наблюдения и результаты
    • Красные — ошибки
  3. Элементы оценки шага рядом с каждой карточкой
  4. Схемы уровня трассы под отображением трассы

Типичный порядок работы:

  1. Прочитайте описание задачи, чтобы понять, что агент должен был сделать
  2. Пройдите по шагам трассы, оценивая каждый
  3. Для шагов, помеченных «частично верно» или «неверно», выберите тип (или типы) ошибки
  4. Оцените трассу целиком (успех, правильность, эффективность)
  5. При необходимости добавьте заметки
  6. Отправьте и переходите к следующей трассе

Советы разметчикам

Разворачивайте свёрнутые наблюдения, а не доверяйте им на глаз: именно там ловится момент, когда агент неверно прочитал найденное. Сверьте итоговый ответ с эталоном, если он есть, прежде чем оценивать успешность задачи. Различайте «лишний» и «неверный» шаг: лишний тратит усилия, но ошибки не вносит. А на длинных трассах боковой таймлайн шагов позволяет сразу перейти к нужному шагу.


Шаг 6: разбор результатов

После разметки результаты разбираются программно.

Базовый разбор через pandas

python
import pandas as pd
import json
 
# Load annotations
annotations = []
with open("output/annotations.jsonl") as f:
    for line in f:
        annotations.append(json.loads(line))
 
df = pd.DataFrame(annotations)
 
# Task success rate
success_counts = df.groupby("annotations").apply(
    lambda x: x.iloc[0]["annotations"]["task_success"]
).value_counts()
print("Task Success Distribution:")
print(success_counts)
 
# Average efficiency rating
efficiency_scores = [
    a["annotations"]["efficiency"]
    for a in annotations
    if "efficiency" in a["annotations"]
]
print(f"\nAverage Efficiency: {sum(efficiency_scores) / len(efficiency_scores):.2f}")

Разбор ошибок по шагам

python
# Collect all step-level errors
error_counts = {}
for ann in annotations:
    step_errors = ann["annotations"].get("error_type", {})
    for step_idx, errors in step_errors.items():
        for error in errors:
            error_counts[error] = error_counts.get(error, 0) + 1
 
print("Error Type Distribution:")
for error, count in sorted(error_counts.items(), key=lambda x: -x[1]):
    print(f"  {error}: {count}")

Разбор через DuckDB (по Parquet)

python
import duckdb
 
# Overall success rate
result = duckdb.sql("""
    SELECT value, COUNT(*) as count
    FROM 'output/parquet/annotations.parquet'
    WHERE schema_name = 'task_success'
    GROUP BY value
    ORDER BY count DESC
""")
print(result)

Шаг 7: масштабирование

Для более крупных проектов оценки (сотни и тысячи трасс) присмотритесь к таким настройкам:

Несколько разметчиков

Назначайте на трассу нескольких разметчиков ради согласованности:

yaml
annotation_task_config:
  total_annotations_per_instance: 3
  assignment_strategy: random

Готовые схемы

Для быстрой настройки возьмите готовые схемы оценки агентов из Potato:

yaml
annotation_schemes:
  - preset: agent_task_success
  - preset: agent_step_correctness
  - preset: agent_error_taxonomy
  - preset: agent_efficiency

Контроль качества

Включите эталонные объекты для наблюдения за качеством:

yaml
phases:
  training:
    enabled: true
    data_file: "data/training_traces.jsonl"
    passing_criteria:
      min_correct: 4
      total_questions: 5

Как приспособить под другие типы агентов

Вызов функций OpenAI

yaml
agentic:
  enabled: true
  trace_converter: openai
  display_type: agent_trace

Работа с инструментами Anthropic

yaml
agentic:
  enabled: true
  trace_converter: anthropic
  display_type: agent_trace

Мультиагентные системы (CrewAI/AutoGen)

yaml
agentic:
  enabled: true
  trace_converter: multi_agent
  display_type: agent_trace
  multi_agent:
    agent_converters:
      researcher: react
      writer: anthropic
      reviewer: openai

Для кодовых агентов Potato отрисовывает диффы кода и вывод терминала с нормальным форматированием:

Coding agent evaluation in Potato

Агенты, работающие в вебе

Для веб-агентов переключитесь на отображение веб-агента:

yaml
agentic:
  enabled: true
  trace_converter: webarena
  display_type: web_agent
  web_agent_display:
    screenshot_max_width: 900
    overlay:
      enabled: true
    filmstrip:
      enabled: true

Отдельное руководство — Разметка агентов, работающих в вебе.


Итог

Система разметки агентов в Potato идёт с 15 конвертерами трасс, так что трассы из большинства фреймворков заводятся без переделки, с тремя типами отображения под работу с инструментами, веб-сёрфинг и диалоговых агентов, с оценками по репликам для суждений на уровне шага, с 9 готовыми схемами под частые измерения оценки и с экспортом в Parquet — для момента, когда вы садитесь всё это разбирать.

«Дал ли агент правильный ответ?» — простой вопрос. Верно ли он рассуждал на каждом шаге — вопрос потруднее, и отвечает на него именно пошаговая разметка. Агрегированные метрики эти ошибки усредняют и прячут.

Детали реализации — в руководстве по оценке агентов и в документации по трассам агентов.


Что почитать дальше