代码评审标注
用 GitHub PR 风格的内联 diff 评论、文件级正确性评分,以及通过或打回的结论,评审 AI 编码智能体的输出,衡量代码质量。
v2.4.0 新增
评估 AI 编码智能体产出的代码改动,光靠一个通过/失败的判断远远不够。研究者和工程团队需要在多个粒度上衡量代码质量:某一行可能藏着 bug 或风格问题,某个文件可能改对了也可能根本不该改,而整套改动可能解决了问题却留下了技术债。这正是人类评审在 GitHub 上看 pull request 时走的流程。
Potato 的代码评审标注模式把 GitHub PR 评审的体验带进了智能体评估。标注者能看到智能体改动过的每个文件的 unified diff。他们可以点击任意 diff 行,留下带分类标签的内联评论。每个文件都会得到一个正确性和质量评分。标注者最后给出结论:approve、request changes 或 comment only。全部内容都被记录为结构化的标注数据,可直接用于训练代码质量模型。
内联评论
标注者点击 diff 中的任意一行,打开内联评论表单。每条评论包含一个分类、一个严重程度和自由文本内容。评论会锚定在具体的那一行上,和 GitHub PR 评审评论一样。
评论分类
默认的评论分类覆盖了最常见的代码评审反馈类型:
| 分类 | 说明 |
|---|---|
bug | 功能性 bug —— 代码无法正确工作 |
logic | 逻辑错误 —— 语法没问题,但做法本身有缺陷 |
security | 安全漏洞或不安全的做法 |
performance | 性能问题 —— 多余的计算、内存泄漏等 |
style | 风格问题 —— 命名、格式、是否地道 |
suggestion | 有更好的替代做法 |
question | 需要澄清 —— 评审者拿不准这里的意图 |
praise | 正面反馈 —— 智能体做得好的地方 |
配置
annotation_schemes:
- annotation_type: code_review
name: review
description: "Click any diff line to add an inline comment"
# Categories offered on each inline comment
comment_categories:
- bug
- logic
- security
- performance
- style
- suggestion
- question建议的代码改动
启用 allow_suggestions 后,标注者可以为自己正在评论的代码块写一段替换建议,对应 GitHub 的 “suggestion” 功能。建议会以代码块的形式出现在评论下方,可用于训练代码修复模型。
{
"category": "bug",
"file": "src/parser.py",
"line": 42,
"text": "Off-by-one error: range should be inclusive of end"
}文件级评分
智能体改动过的每个文件都会得到两项彼此独立的评分:正确性和代码质量。
配置
annotation_schemes:
- annotation_type: code_review
name: review
description: "Rate each modified file"
# One 1-5 rating per dimension, per file touched by the diff
file_rating_dimensions:
- correctness
- quality输出格式
{
"file_ratings": {
"src/parser.py": {
"correctness": 4,
"quality": 3
},
"tests/test_parser.py": {
"correctness": 5,
"quality": 4
},
"src/utils.py": {
"correctness": 2,
"quality": 2
}
}
}总体结论
看完所有文件、留下内联评论之后,标注者为整套改动给出一个总体结论。
配置
annotation_schemes:
- annotation_type: code_review
name: review
description: "Give an overall verdict on the code changes"
verdict_options:
- approve
- request_changes
- comment_only配置参考
一个代码评审标注任务的完整配置:
annotation_task_name: "Coding Agent Code Review"
task_dir: "."
data_files:
- "data/coding_traces.jsonl"
item_properties:
id_key: id
text_key: task_description
instance_display:
fields:
- key: structured_turns
type: coding_trace
label: "Agent changes"
display_options:
diff_view: unified
terminal_theme: dark
collapse_long_outputs: true
max_output_lines: 50
show_file_tree: true
show_step_numbers: true
show_reasoning: true
annotation_schemes:
# Inline comments, file ratings and the overall verdict are all one scheme
- annotation_type: code_review
name: review
description: "Review the agent's code changes"
comment_categories:
- bug
- logic
- security
- performance
- style
- suggestion
- question
- praise
file_rating_dimensions:
- correctness
- quality
verdict_options:
- approve
- request_changes
- comment_only
# A free-text summary is a separate scheme
- annotation_type: text
name: summary
description: "Summarize your review"
rows: 4
output_annotation_dir: "output/"
export_annotation_format: "jsonl"标注流程
标注者在完成一个代码评审标注任务时,看到和要做的是这些:
-
任务概览:任务描述显示在最上方,说明智能体被要求做什么(比如“修复 test_parser.py 中失败的测试”)。
-
文件树导航:左侧栏列出智能体动过的所有文件,并用颜色区分:绿色是新增文件,黄色是修改的文件,红色是删除的文件。
-
审阅 diff:主面板按文件显示 unified diff。标注者逐段翻看,读每一处改动。
-
添加内联评论:点击行号会打开评论表单。标注者选择一个分类(bug、建议等),可选地选择严重程度,写下评论,也可以附上一段代码建议。
-
文件评分:看完某个文件的 diff 之后,标注者用该文件 diff 下方的评分组件,从正确性(1-5)和代码质量(1-5)两方面打分。
-
总体结论:在页面底部,标注者选择一个结论(通过、要求修改,或只留评论),并写一段评审总结。
-
提交:标注者点击 “Submit”,把全部内联评论、文件评分和结论作为一条标注记录保存。
数据格式
一次代码评审标注的完整输出:
{
"instance_id": "trace_042",
"annotator": "reviewer_01",
"verdict": "request_changes",
"comments": [
{
"category": "bug",
"file": "src/parser.py",
"line": 42,
"text": "This will throw IndexError when tokens list is empty"
},
{
"category": "style",
"file": "src/parser.py",
"line": 15,
"text": "Variable name 'x' is not descriptive"
},
{
"category": "praise",
"file": "tests/test_parser.py",
"line": 28,
"text": "Good edge case coverage for empty input"
}
],
"file_ratings": {
"src/parser.py": { "correctness": 3, "quality": 2 },
"tests/test_parser.py": { "correctness": 5, "quality": 4 }
}
}导出
代码评审标注通过 coding_eval 导出器产出。它接受的是配置文件而不是输出目录,并写入你指定的目录:
python -m potato.export \
-c config.yaml \
-f coding_eval \
-o results/ \
--option types=code_reviewtypes 决定写出哪些内容,可取 prm、preference、swebench 和 code_review,默认四种全写。去掉 --option 就能在评审之外一并拿到其余几种。如果某个请求的类型在这项研究里没有数据,导出器会明说,而不是写出一个空文件。
代码评审的输出里带着每条内联评论及其文件、行号、分类和严重程度,还有逐文件的评分和每次评审的结论,所以评论级的训练数据和评分分析都来自同一个文件。
另请参阅
- 编码智能体标注 —— 展示编码智能体 trace,带 diff 渲染和文件树
- 过程奖励标注 —— 为 PRM 训练收集逐步奖励信号
- 实时编码智能体观察 —— 实时观察编码智能体并与之交互
- 智能体标注 —— 通用的智能体 trace 标注
- 导出格式 —— 所有支持的导出格式
有关实现详情,请参阅源文档。