Skip to content

代码评审标注

用 GitHub PR 风格的内联 diff 评论、文件级正确性评分,以及通过或打回的结论,评审 AI 编码智能体的输出,衡量代码质量。

v2.4.0 新增

评估 AI 编码智能体产出的代码改动,光靠一个通过/失败的判断远远不够。研究者和工程团队需要在多个粒度上衡量代码质量:某一行可能藏着 bug 或风格问题,某个文件可能改对了也可能根本不该改,而整套改动可能解决了问题却留下了技术债。这正是人类评审在 GitHub 上看 pull request 时走的流程。

Potato 的代码评审标注模式把 GitHub PR 评审的体验带进了智能体评估。标注者能看到智能体改动过的每个文件的 unified diff。他们可以点击任意 diff 行,留下带分类标签的内联评论。每个文件都会得到一个正确性和质量评分。标注者最后给出结论:approve、request changes 或 comment only。全部内容都被记录为结构化的标注数据,可直接用于训练代码质量模型。

内联评论

标注者点击 diff 中的任意一行,打开内联评论表单。每条评论包含一个分类、一个严重程度和自由文本内容。评论会锚定在具体的那一行上,和 GitHub PR 评审评论一样。

评论分类

默认的评论分类覆盖了最常见的代码评审反馈类型:

分类说明
bug功能性 bug —— 代码无法正确工作
logic逻辑错误 —— 语法没问题,但做法本身有缺陷
security安全漏洞或不安全的做法
performance性能问题 —— 多余的计算、内存泄漏等
style风格问题 —— 命名、格式、是否地道
suggestion有更好的替代做法
question需要澄清 —— 评审者拿不准这里的意图
praise正面反馈 —— 智能体做得好的地方

配置

yaml
annotation_schemes:
  - annotation_type: code_review
    name: review
    description: "Click any diff line to add an inline comment"
 
    # Categories offered on each inline comment
    comment_categories:
      - bug
      - logic
      - security
      - performance
      - style
      - suggestion
      - question

建议的代码改动

启用 allow_suggestions 后,标注者可以为自己正在评论的代码块写一段替换建议,对应 GitHub 的 “suggestion” 功能。建议会以代码块的形式出现在评论下方,可用于训练代码修复模型。

json
{
  "category": "bug",
  "file": "src/parser.py",
  "line": 42,
  "text": "Off-by-one error: range should be inclusive of end"
}

文件级评分

智能体改动过的每个文件都会得到两项彼此独立的评分:正确性和代码质量。

配置

yaml
annotation_schemes:
  - annotation_type: code_review
    name: review
    description: "Rate each modified file"
 
    # One 1-5 rating per dimension, per file touched by the diff
    file_rating_dimensions:
      - correctness
      - quality

输出格式

json
{
  "file_ratings": {
    "src/parser.py": {
      "correctness": 4,
      "quality": 3
    },
    "tests/test_parser.py": {
      "correctness": 5,
      "quality": 4
    },
    "src/utils.py": {
      "correctness": 2,
      "quality": 2
    }
  }
}

总体结论

看完所有文件、留下内联评论之后,标注者为整套改动给出一个总体结论。

配置

yaml
annotation_schemes:
  - annotation_type: code_review
    name: review
    description: "Give an overall verdict on the code changes"
 
    verdict_options:
      - approve
      - request_changes
      - comment_only

配置参考

一个代码评审标注任务的完整配置:

yaml
annotation_task_name: "Coding Agent Code Review"
task_dir: "."
 
data_files:
  - "data/coding_traces.jsonl"
 
item_properties:
  id_key: id
  text_key: task_description
 
instance_display:
  fields:
    - key: structured_turns
      type: coding_trace
      label: "Agent changes"
      display_options:
        diff_view: unified
        terminal_theme: dark
        collapse_long_outputs: true
        max_output_lines: 50
        show_file_tree: true
        show_step_numbers: true
        show_reasoning: true
 
annotation_schemes:
  # Inline comments, file ratings and the overall verdict are all one scheme
  - annotation_type: code_review
    name: review
    description: "Review the agent's code changes"
    comment_categories:
      - bug
      - logic
      - security
      - performance
      - style
      - suggestion
      - question
      - praise
    file_rating_dimensions:
      - correctness
      - quality
    verdict_options:
      - approve
      - request_changes
      - comment_only
 
  # A free-text summary is a separate scheme
  - annotation_type: text
    name: summary
    description: "Summarize your review"
    rows: 4
 
output_annotation_dir: "output/"
export_annotation_format: "jsonl"

标注流程

标注者在完成一个代码评审标注任务时,看到和要做的是这些:

  1. 任务概览:任务描述显示在最上方,说明智能体被要求做什么(比如“修复 test_parser.py 中失败的测试”)。

  2. 文件树导航:左侧栏列出智能体动过的所有文件,并用颜色区分:绿色是新增文件,黄色是修改的文件,红色是删除的文件。

  3. 审阅 diff:主面板按文件显示 unified diff。标注者逐段翻看,读每一处改动。

  4. 添加内联评论:点击行号会打开评论表单。标注者选择一个分类(bug、建议等),可选地选择严重程度,写下评论,也可以附上一段代码建议。

  5. 文件评分:看完某个文件的 diff 之后,标注者用该文件 diff 下方的评分组件,从正确性(1-5)和代码质量(1-5)两方面打分。

  6. 总体结论:在页面底部,标注者选择一个结论(通过、要求修改,或只留评论),并写一段评审总结。

  7. 提交:标注者点击 “Submit”,把全部内联评论、文件评分和结论作为一条标注记录保存。

数据格式

一次代码评审标注的完整输出:

json
{
  "instance_id": "trace_042",
  "annotator": "reviewer_01",
  "verdict": "request_changes",
  "comments": [
    {
      "category": "bug",
      "file": "src/parser.py",
      "line": 42,
      "text": "This will throw IndexError when tokens list is empty"
    },
    {
      "category": "style",
      "file": "src/parser.py",
      "line": 15,
      "text": "Variable name 'x' is not descriptive"
    },
    {
      "category": "praise",
      "file": "tests/test_parser.py",
      "line": 28,
      "text": "Good edge case coverage for empty input"
    }
  ],
  "file_ratings": {
    "src/parser.py": { "correctness": 3, "quality": 2 },
    "tests/test_parser.py": { "correctness": 5, "quality": 4 }
  }
}

导出

代码评审标注通过 coding_eval 导出器产出。它接受的是配置文件而不是输出目录,并写入你指定的目录:

bash
python -m potato.export \
  -c config.yaml \
  -f coding_eval \
  -o results/ \
  --option types=code_review

types 决定写出哪些内容,可取 prm、preference、swebench 和 code_review,默认四种全写。去掉 --option 就能在评审之外一并拿到其余几种。如果某个请求的类型在这项研究里没有数据,导出器会明说,而不是写出一个空文件。

代码评审的输出里带着每条内联评论及其文件、行号、分类和严重程度,还有逐文件的评分和每次评审的结论,所以评论级的训练数据和评分分析都来自同一个文件。

另请参阅

有关实现详情,请参阅源文档。