Skip to content

过程奖励标注

用首错模式和逐步模式为过程奖励模型收集每一步的奖励标签,可选中性标签,并支持由人工核验的 LLM 预标注。

v2.4.0 新增

过程奖励模型评的是模型推理过程中的每一步,而不只是最终答案,所以训练它需要的是每步一个标签,而不是每条轨迹一个标签。process_reward 方案负责收集这些标签。

两种模式在速度和细致程度之间做取舍。首错模式只要一次点击:第一个出错的步骤,它之后的一切都视为受牵连。逐步模式要求对每一步都做判断,代价更大,但能捕捉首错模式抓不到的情况,比如智能体从自己的错误里恢复过来。

两种模式都能配合编码轨迹、智能体轨迹和思维链这几种显示方式使用。

首错模式

标注者顺着轨迹往下读,点中第一个不正确的步骤。它之前的步骤记为正确,被点中的这一步以及之后的全部记为不正确。这样产生的标签序列正是二元 PRM 训练所需要的:错误之前是 1,从错误开始是 -1。

yaml
annotation_schemes:
  - annotation_type: process_reward
    name: step_rewards
    description: "Click the first step where the agent made a mistake"
    mode: first_error
    steps_key: structured_turns
    step_text_key: content

steps_key 指定存放步骤列表的字段,step_text_key 指定每个步骤里用来显示的字段。

标注流程

  1. 步骤起初都未标记
  2. 标注者点击第一个不正确的步骤
  3. 它之前的一切变为绿色
  4. 该步骤被标记为首个错误
  5. 它之后的一切被标记为受牵连

颜色把这种连带关系显示了出来,而这正是这个模式的用意:一次判断就标完整条轨迹。

逐步模式

每一步都单独判断。加上 allow_neutral 可以得到 PRM800K 那种三分类标签,其中一步可以既不算对也不算错:

yaml
annotation_schemes:
  - annotation_type: process_reward
    name: step_rewards
    description: "Rate each step correct, neutral, or incorrect"
    mode: per_step
    steps_key: structured_turns
    step_text_key: content
    allow_neutral: true
    inline_with_trace: true

allow_neutral 只在逐步模式下起作用。首错模式的连带关系里没有中性步骤的位置,所以这个选项在那里会被忽略。

inline_with_trace: true 把评分控件挂到显示中的每一步上,而不是收在单独的面板里。轨迹很长时你会想要这个,否则标注者容易找不着自己看到哪了。

给标签改名

三个按钮默认显示 Correct、Neutral 和 Wrong。在这些词不合适的领域里,reward_labels 可以给它们改名:

yaml
annotation_schemes:
  - annotation_type: process_reward
    name: step_rewards
    description: "Rate each reasoning step"
    mode: per_step
    steps_key: cot_steps
    allow_neutral: true
    reward_labels:
      correct: "Valid"
      neutral: "Neutral"
      incorrect: "Flawed"

切分思维链

推理轨迹通常是一整条长字符串,而不是步骤列表。cot_segmentation 会在标注开始前把它切开:

yaml
cot_segmentation:
  source_key: reasoning      # the item field holding the long CoT string
  strategy: auto             # blank_line | numbered | markers | sentence | llm | auto
  target_key: cot_steps      # where the step list is written
  min_step_chars: 30         # merge anything shorter into the previous step
  max_steps: 200
 
instance_display:
  fields:
    - key: cot_steps
      type: cot_trace
      label: "Reasoning"
 
annotation_schemes:
  - annotation_type: process_reward
    name: step_rewards
    description: "Rate each reasoning step"
    mode: per_step
    steps_key: cot_steps
    allow_neutral: true
    inline_with_trace: true

min_step_chars 值得设一下。不设的话,按句子和按空行切分会产生一行的碎片,标注者孤立地看根本没法判断,拿回来的标签就是噪声。

LLM 预标注

可以让模型先给出每一步的标签,标注者再逐个确认或改掉,在长轨迹上这比从零开始打标签要快:

yaml
annotation_schemes:
  - annotation_type: process_reward
    name: step_rewards
    description: "Verify or correct each suggested reward"
    mode: per_step
    steps_key: cot_steps
    allow_neutral: true
    inline_with_trace: true
    ai_prelabel: true
    require_verification: true
 
ai_support:
  enabled: true
  endpoint_type: openai
  ai_config:
    model: gpt-4o-mini
    max_tokens: 768
    temperature: 0.1
    include:
      all: true

ai_prelabel 加上那个请求建议的按钮。require_verification 会在每条建议都被确认或修改之前拒绝接受这份标注,正是它保证了输出是人给的标签,而不是模型给的。

max_tokens 要留足余量。一条长思维链加上每一步的判定是一份很大的回复,而被截断的回复呈现出来就是什么建议都没有。

输出里有什么

标注以方案名为键存放,内容是一个 steps 列表,外加收集时所用的 mode:

json
{
  "step_rewards": {
    "mode": "per_step",
    "steps": [
      {"reward": 1, "source": "human", "verified": true},
      {"reward": -1, "source": "human", "verified": true},
      {"reward": 0, "source": "ai", "verified": true, "ai_reward": 0},
      {"reward": null, "source": null, "verified": false}
    ]
  }
}

reward 为 1 表示正确,-1 表示不正确,0 表示中性。被跳过的步骤是 null 而不是 0,所以没人判断过的步骤和被判为既不对也不错的步骤仍然分得清。这个区别在标签变成训练数据时很重要,因为把跳过当成中性等于凭空造出信号。

source 记录这个值来自人还是来自模型,verified 记录是否有人确认过。开着 ai_prelabel 时,即使标注者改掉了建议,ai_reward 也会保留模型最初的那个,正是它让你不必再跑一遍就能报告模型对的频率。

导出

Potato 没有内置的 PRM 或 DPO 导出格式。把标注导出来,自己整理成需要的形状:

bash
python -m potato.export -c config.yaml -f jsonl -o ./export/

把首错标注转成步骤级的训练标签是段很短的代码,而自己来写,才能由你决定跳过的步骤怎么处理:

python
import json
from pathlib import Path
 
SCHEME = "step_rewards"
 
examples = []
for f in Path("export/").rglob("*.jsonl"):
    with open(f) as fh:
        for line in fh:
            rec = json.loads(line)
            payload = (rec.get("labels", rec) or {}).get(SCHEME)
            if not payload:
                continue
 
            steps = payload.get("steps", [])
            # Drop traces with unjudged steps rather than guessing at them.
            if any(s.get("reward") is None for s in steps):
                continue
 
            item = items_by_id[rec["instance_id"]]      # your own data file
            examples.append({
                "trace_id": rec["instance_id"],
                "steps": [
                    {"content": src, "label": s["reward"]}
                    for src, s in zip(item["structured_turns"], steps)
                ],
            })
 
with open("prm_training_data.jsonl", "w") as out:
    for ex in examples:
        out.write(json.dumps(ex) + "\n")
 
print(f"Wrote {len(examples)} traces")

分析

python
import json
from collections import Counter
from pathlib import Path
 
SCHEME = "step_rewards"
 
traces = []
for f in Path("export/").rglob("*.jsonl"):
    with open(f) as fh:
        for line in fh:
            rec = json.loads(line)
            payload = (rec.get("labels", rec) or {}).get(SCHEME)
            if payload:
                traces.append(payload)
 
print(f"Traces annotated: {len(traces)}")
 
all_correct = sum(
    1 for t in traces
    if all(s.get("reward") == 1 for s in t.get("steps", []))
)
print(f"Traces with no errors: {all_correct} ({all_correct / len(traces):.1%})")
 
# Where the first error lands
first_error_positions = []
for t in traces:
    for i, s in enumerate(t.get("steps", [])):
        if s.get("reward") == -1:
            first_error_positions.append(i)
            break
 
if first_error_positions:
    avg = sum(first_error_positions) / len(first_error_positions)
    print(f"Average first-error position: step {avg:.1f}")
    print("By position:", Counter(first_error_positions).most_common(10))

开着 ai_prelabel 时,同一批记录还能告诉你模型的建议有多少经受住了审核:

python
kept = changed = 0
for t in traces:
    for s in t.get("steps", []):
        if s.get("ai_reward") is None:
            continue
        if s.get("reward") == s.get("ai_reward"):
            kept += 1
        else:
            changed += 1
 
total = kept + changed
if total:
    print(f"AI suggestions kept: {kept}/{total} ({kept / total:.1%})")

在一个新领域里信任预标注之前,这个接受率就是要盯的数字。接近 100% 的比率通常意味着标注者是在确认而不是在检查,而不是模型真有那么准。

研究背景

过程奖励标注的存在是为了支撑多步系统上的奖励模型研究,推动整套做法的那个发现是:正确的最终答案可以架在有缺陷的推理之上。步骤级的标签把两者分开,而这里的首错模式和逐步模式对应的正是这一脉研究所用的标签格式。

关于这套分类体系的设计以及随之而来的一致性问题,见逐步错误定位,它从错误分类的角度讲同样的轨迹。

另请参阅

实现细节见源文档。