Skip to content

コードレビューアノテーション

GitHub PRスタイルのインラインdiffコメント、ファイル単位の正確性評価、承認・却下の判定を使って、AIコーディングエージェントの出力をコード品質の観点からレビューします。

v2.4.0の新機能

AIコーディングエージェントが生成したコード変更の評価には、合否の二値判断以上のものが要ります。研究者やエンジニアリングチームは、複数の粒度でコード品質を見る必要があります。個々の行にバグやスタイル違反が混ざることもあれば、ファイル単位で見て変更が適切な場合も不要な場合もあり、変更セット全体としては問題を解いていても技術的負債を持ち込んでいることもあります。これは、GitHubでpull requestをレビューするときに人間のレビュアーが辿るのと同じ流れです。

Potatoのコードレビューアノテーションモードは、GitHubのPRレビュー体験をエージェント評価に持ち込みます。アノテーターは、エージェントが変更したすべてのファイルについて統一diffを見ます。diffの任意の行をクリックして、カテゴリタグ付きのインラインコメントを残せます。各ファイルには正確性と品質の評価が付きます。最後にアノテーターが判定を下します。approve、request changes、comment onlyのいずれかです。これらはすべて構造化されたアノテーションデータとして記録され、コード品質モデルの訓練にそのまま使えます。

インラインコメント

アノテーターはdiffの任意の行をクリックして、インラインコメントのフォームを開きます。各コメントにはカテゴリ、重大度、自由記述の内容があります。コメントはGitHubのPRレビューコメントと同じように、特定の行に紐づいて表示されます。

コメントカテゴリ

デフォルトのコメントカテゴリは、コードレビューで最もよく使われるフィードバックの種類を網羅しています。

カテゴリ説明
bug機能上のバグ -- コードが正しく動作しない
logicロジックの誤り -- 構文は正しくてもアプローチに欠陥がある
securityセキュリティ脆弱性、または安全でない書き方
performance性能上の問題 -- 不要な計算、メモリリークなど
styleスタイル違反 -- 命名、書式、イディオムの使い方
suggestionより良い代替アプローチ
question確認が必要 -- レビュアーが意図を掴めていない
praise肯定的なフィードバック -- エージェントがうまくやった点

設定

yaml
annotation_schemes:
  - annotation_type: code_review
    name: review
    description: "Click any diff line to add an inline comment"
 
    # Categories offered on each inline comment
    comment_categories:
      - bug
      - logic
      - security
      - performance
      - style
      - suggestion
      - question

コードの修正提案

allow_suggestionsを有効にすると、アノテーターはコメント対象のコードブロックに対する置き換え案を書けます。GitHubの「suggestion」機能に相当します。提案はコメントの下にコードブロックとして表示され、コード修復モデルの訓練に利用できます。

json
{
  "category": "bug",
  "file": "src/parser.py",
  "line": 42,
  "text": "Off-by-one error: range should be inclusive of end"
}

ファイル単位の評価

エージェントが変更した各ファイルには、正確性とコード品質という2つの独立した評価が付きます。

設定

yaml
annotation_schemes:
  - annotation_type: code_review
    name: review
    description: "Rate each modified file"
 
    # One 1-5 rating per dimension, per file touched by the diff
    file_rating_dimensions:
      - correctness
      - quality

出力形式

json
{
  "file_ratings": {
    "src/parser.py": {
      "correctness": 4,
      "quality": 3
    },
    "tests/test_parser.py": {
      "correctness": 5,
      "quality": 4
    },
    "src/utils.py": {
      "correctness": 2,
      "quality": 2
    }
  }
}

全体の判定

すべてのファイルをレビューしてインラインコメントを残したあと、アノテーターは変更セット全体に対する判定を下します。

設定

yaml
annotation_schemes:
  - annotation_type: code_review
    name: review
    description: "Give an overall verdict on the code changes"
 
    verdict_options:
      - approve
      - request_changes
      - comment_only

設定リファレンス

コードレビューアノテーションタスクの完全な設定です。

yaml
annotation_task_name: "Coding Agent Code Review"
task_dir: "."
 
data_files:
  - "data/coding_traces.jsonl"
 
item_properties:
  id_key: id
  text_key: task_description
 
instance_display:
  fields:
    - key: structured_turns
      type: coding_trace
      label: "Agent changes"
      display_options:
        diff_view: unified
        terminal_theme: dark
        collapse_long_outputs: true
        max_output_lines: 50
        show_file_tree: true
        show_step_numbers: true
        show_reasoning: true
 
annotation_schemes:
  # Inline comments, file ratings and the overall verdict are all one scheme
  - annotation_type: code_review
    name: review
    description: "Review the agent's code changes"
    comment_categories:
      - bug
      - logic
      - security
      - performance
      - style
      - suggestion
      - question
      - praise
    file_rating_dimensions:
      - correctness
      - quality
    verdict_options:
      - approve
      - request_changes
      - comment_only
 
  # A free-text summary is a separate scheme
  - annotation_type: text
    name: summary
    description: "Summarize your review"
    rows: 4
 
output_annotation_dir: "output/"
export_annotation_format: "jsonl"

アノテーションの流れ

コードレビューアノテーションタスクで、アノテーターが目にするものと行うことは次のとおりです。

  1. タスクの概要:上部にタスクの説明が表示され、エージェントが何を指示されたかが分かります(例:「test_parser.pyの失敗しているテストを直す」)。

  2. ファイルツリーの操作:左サイドバーにエージェントが触れたすべてのファイルが並びます。ファイルは色分けされていて、新規は緑、変更は黄、削除は赤です。

  3. diffのレビュー:メインパネルに各ファイルの統一diffが表示されます。アノテーターはdiffをスクロールしながら、変更を1つずつ読んでいきます。

  4. インラインコメントの追加:行番号をクリックするとコメントフォームが開きます。アノテーターはカテゴリ(bug、suggestionなど)を選び、必要に応じて重大度を選び、コメントを書き、必要ならコードの修正提案を添えます。

  5. ファイルの評価:各ファイルのdiffをレビューしたあと、diffの下にある評価ウィジェットで正確性(1〜5)とコード品質(1〜5)を付けます。

  6. 全体の判定:一番下で判定(approve、request changes、comment only)を選び、レビューの要約を書きます。

  7. 送信:「Submit」をクリックすると、インラインコメント、ファイル評価、判定が1件のアノテーションレコードとして保存されます。

データ形式

1件のコードレビューアノテーションの完全な出力です。

json
{
  "instance_id": "trace_042",
  "annotator": "reviewer_01",
  "verdict": "request_changes",
  "comments": [
    {
      "category": "bug",
      "file": "src/parser.py",
      "line": 42,
      "text": "This will throw IndexError when tokens list is empty"
    },
    {
      "category": "style",
      "file": "src/parser.py",
      "line": 15,
      "text": "Variable name 'x' is not descriptive"
    },
    {
      "category": "praise",
      "file": "tests/test_parser.py",
      "line": 28,
      "text": "Good edge case coverage for empty input"
    }
  ],
  "file_ratings": {
    "src/parser.py": { "correctness": 3, "quality": 2 },
    "tests/test_parser.py": { "correctness": 5, "quality": 4 }
  }
}

エクスポート

コードレビューのアノテーションは、いくつかの形式でエクスポートできます。

bash
python -m potato.export \
  -c config.yaml \
  -f coding_eval \
  -o results/ \
  --option types=code_review

code_review_comments形式は、コードレビューコメントを生成するモデルや、コードの問題箇所とそのカテゴリを予測するモデルの訓練に特に向いています。

参考資料

実装の詳細については、ソースドキュメントを参照してください。