Back

CodeHawk

AI-powered pull request review pipeline

An AI reviewer that reads the pull request before the lead engineer does.

Stack

PythonOpenAI AgentsDockerAzure DevOpsGitHub

Python · MIT

Problem

  • Last approval always landed on one lead engineer
  • PRs sat until an hour opened up on the calendar
  • AI-assisted changesets were getting much larger
  • Team conventions slipping when people were in a hurry
  • A 3-line helper change looked like a 3-line fixture
  • Scale of impact was invisible to the reviewer

Solution

  • Files risk-scored, batched ten at a time, three concurrent
  • 16 language rule sets and 6 review modes
  • Two-phase: agent writes findings.json, engine posts
  • AST graph for impact radius and missing tests
  • Scan cheap, then verify expensive with tools
  • Docker image for Azure DevOps and GitHub CI

Outcome

  • Lead-engineer bottleneck off the critical path
  • Inline comments posted without a human in the loop
  • Configurable star rating and CI gate
  • Re-push verifies fixes without a full re-review
  • Dry-run and replay from the same findings.json
  • Open-sourced after solving it on a team of ten

16

language rule sets

6

review modes

2-phase

agent then poster

394

unit tests

Architecture Review

System design · Pipeline · Decisions

System Architecture

How It Works

  1. 1

    Prepare

    CI runs the container. Before any model call: fetch the PR, drop non-code files, check existing threads, build the AST graph if it can, risk-score every file, and split into batches of 10.

  2. 2

    Scan

    Each batch gets two single-turn calls with the diffs already in the prompt. Pass 1A covers correctness and testing; Pass 1B covers security, performance, and architecture. Neither has tools. Candidates merge and dedupe on file, line, and category.

  3. 3

    Verify

    A short agent loop picks up the candidates and now gets tools: file reads, ripgrep, git blame, graph queries. It confirms or drops each finding. If this pass falls over, Pass 1 candidates still ship at lower confidence.

  4. 4

    Score and post

    Phase 2 validates against the schema, applies a confidence floor, then scores with a penalty matrix mapped to 0–5 stars. Comments go inline, a summary lands on the PR, and CI gets a structured pass/fail.

  5. 5

    Re-push

    Deleted-file findings drop, untouched files stay open with no model call, and only modified files get re-verified. A developer can also reply and argue — if it holds, the thread resolves as WONT_FIX with a suggested .codereview.md rule.

Key Decisions

The agent writes data. It never posts.

Phase 1 produces findings.json and stops. Phase 2 does everything with consequences: validation, scoring, comments, the CI gate. Phase 2 is testable without spending a token, --dry-run is a real path, and a botched post can be replayed without paying for the review twice.

Scan cheap, then verify expensive

A single long agent loop on a 98-file PR burned 11.3M tokens. Two single-turn calls now do the reasoning against diffs already in the prompt; a short tool-using loop verifies candidates and throws out the false ones.

The graph decides how much of the PR actually gets read

An AST graph of CALLS and IMPORTS_FROM feeds a per-file risk score. HIGH files get a full read, MEDIUM get the diff plus a read if needed, LOW get a scan. The agent accounts for every file, but not with the same attention.

Split the work, don't summarise it

When a diff won't fit, compressing it throws away the detail worth reviewing. Files batch ten at a time instead. Only a single file that is still too big falls back to hunk summaries, and the agent can drill back in at full fidelity.

The poster computes the dedup ID, not the model

Findings get cr-id: sha1(file:line:category), embedded as an HTML comment in the thread. The agent writes cr_id: null; Phase 2 fills it in. Models cannot reliably compute a hash, and a near-miss posts two comments on the same line.

CodeHawk — Architecture

CodeHawk architecture

Evaluation

No labelled benchmark — that's the honest gap. There's no held-out set of PRs with known bugs, so I can't quote precision or recall. What exists: 394 unit tests covering deterministic paths and ugly LLM failure modes, plus dogfooding — CodeHawk reviews its own pull requests. Token cost got measured properly because 11.3M tokens on one PR is a number you notice. A labelled set of about 50 PRs is the first thing I'd build if I picked this up again.

What I'd Change

  • Build the eval harness before the second prompt revision
  • A prompt that references a file is not a prompt that contains it
  • The token optimisation was the token problem
  • cr-id shouldn't have been keyed on the file path

Get in Touch

Let's build
something great.