The agent writes data. It never posts.
Phase 1 produces findings.json and stops. Phase 2 does everything with consequences: validation, scoring, comments, the CI gate. Phase 2 is testable without spending a token, --dry-run is a real path, and a botched post can be replayed without paying for the review twice.
Scan cheap, then verify expensive
A single long agent loop on a 98-file PR burned 11.3M tokens. Two single-turn calls now do the reasoning against diffs already in the prompt; a short tool-using loop verifies candidates and throws out the false ones.
The graph decides how much of the PR actually gets read
An AST graph of CALLS and IMPORTS_FROM feeds a per-file risk score. HIGH files get a full read, MEDIUM get the diff plus a read if needed, LOW get a scan. The agent accounts for every file, but not with the same attention.
Split the work, don't summarise it
When a diff won't fit, compressing it throws away the detail worth reviewing. Files batch ten at a time instead. Only a single file that is still too big falls back to hunk summaries, and the agent can drill back in at full fidelity.
The poster computes the dedup ID, not the model
Findings get cr-id: sha1(file:line:category), embedded as an HTML comment in the thread. The agent writes cr_id: null; Phase 2 fills it in. Models cannot reliably compute a hash, and a near-miss posts two comments on the same line.