# AI Code Review Pipeline

> CI/CD-integrated automated reviews

Automated code reviews that catch patterns humans miss, across 10+ languages.

- **Page:** https://siddharthdeshpande.com/projects/code-review
- **Built at:** Apra Labs
- **Tech stack:** Python, Azure DevOps, GPT, REST APIs

## Impact

- **10+** — languages supported
- **Inline** — PR comments
- **Auto** — fix verification
- **★** — penalty-based scoring

## Problem

- Pull requests waiting hours for review
- Inconsistent review quality across the team
- Recurring patterns missed across pull requests
- No automated quality baseline established
- Senior engineers bottlenecked on reviews
- Style and logic issues caught too late

## Solution

- CI/CD webhook integration with GitHub built
- Penalty-based star rating scoring system
- Fix verification mode for follow-up checks
- 10+ programming languages fully supported
- Contextual inline comments on exact lines
- Configurable rule severity and thresholds

## Outcome

- Automated inline PR comments on every push
- Consistent quality scoring across all repos
- Fix verification closes the feedback loop
- Review bottleneck for seniors eliminated
- Faster merge cycles with fewer regressions
- Quality baseline measurable and tracked

## How it works

1. **Trigger** — Azure DevOps webhook fires on PR events. The system parses the diff and chunks it by file and language for targeted review.
2. **Review** — Language-specific rules combine with GPT analysis to generate inline comments with specific improvement suggestions and a star rating using a penalty-based scoring algorithm.
3. **Verification** — Fix verification mode re-reviews after the developer makes changes, confirming addressed feedback and checking for regressions introduced by the fix.

## Key decisions

- **Penalty-based scoring, not binary pass/fail** — Star ratings with a penalty system for each issue type. Developers get a clear sense of severity and teams can set quality thresholds without blocking every pull request.
- **Fix verification as a separate mode** — After a developer addresses feedback, the system re-reviews just the changed sections. This closes the feedback loop instead of generating a fresh review that might flag new unrelated issues.

## What I'd change

- **Language-specific rules need community input** — Maintaining review rules for 10+ languages is a lot for one team. Open-sourcing the rule engine or integrating community linting configs would scale better.
- **False positive tracking from day one** — No systematic tracking of which review comments developers dismiss. That data would help tune penalty weights and reduce noise over time.
