# Cloud Cost Intelligence Agent

> Automated FinOps analysis

Finding out why the cloud bill moved — every night, before anyone has to ask. Identified 30% savings across the Azure estate.

- **Page:** https://siddharthdeshpande.com/projects/cloud-cost
- **Built at:** Apra Labs
- **Tech stack:** TypeScript, Node 22 / tsx, Azure Cost Management, Azure Resource Graph, Azure Advisor, Azure Monitor, Azure SQL, Claude Sonnet, Azure Service Bus

## Impact

- **30%** — cost reduction
- **19** — data collectors
- **13** — anomaly detectors
- **3** — analysis phases

## Problem

- Cloud spend growing unchecked month over month
- The billing portal shows the total went up, not why
- A schema change can make one query eat a third of database CPU
- Detached disks and idle resources still being paid for months later
- Answering "why" meant stitching cost, telemetry, and inventory by hand
- Quarterly manual audits far too slow — waste compounding silently

## Solution

- 19 collectors across 3 dependency phases — six Azure APIs, SQL DMVs, and the app database
- 13 deterministic anomaly checks and 4-tier breach thresholds
- 20+ standing questions answered from collected data; LLM only for reasoning
- Doer–reviewer loop: an analyst writes, a second agent checks claims against raw files
- Query cost attribution splits the database bill across top queries by CPU share
- Rightsizing, reservation, and unused-resource recommendations, emailed through the existing platform mail path

## Outcome

- 30% cloud cost reduction achieved in months
- Unused resources automatically flagged for cleanup
- Prioritized optimization recommendations delivered nightly
- Teams now own and track their spend
- Anomaly detection over 15-second telemetry samples, reported every night
- Millions saved across the organization annually

## How it works

1. **Collect** — Twelve independent collectors pull cost breakdowns, 30-day trends, resource inventory, orphaned disks and NICs, reservation utilization, Advisor recommendations, seven-day peak metrics, and database telemetry sampled at 15-second intervals. No cross-dependencies, so they parallelize freely. Data is normalized into a common cost model and written as JSON into a per-run directory.
2. **Cross-reference** — Five more collectors read that output back off disk and join it. Query cost attribution splits the database bill across the top queries by CPU share. Breach detection compares each resource against 4-tier configured thresholds. The anomaly detector runs 13 checks: cost spikes above the 30-day average, sustained CPU saturation, a table gaining 500MB overnight, a single query dominating database cost, a reservation about to expire.
3. **Persist** — Breach results and telemetry history land in eight SQL tables. A retention pass trims everything past 365 days on each run.
4. **Analyze** — The deterministic report writer and the LLM loop run in parallel. Arithmetic — day-over-day deltas, threshold breaches, cost per building, which query burned the most CPU — is computed straight from the collected JSON. The LLM gets only the open-ended questions: optimization suggestions, risk assessment, the executive summary. A reviewer agent re-reads the raw files and checks the claims, capped at two rounds. Where both produce an answer, the deterministic one wins.
5. **Deliver** — The merged report goes out by email through the existing platform delivery path — an attachment record and a Service Bus message the platform's mail worker already listens on. Recommendations are prioritized: right-sizing, scheduling, reservation purchases, unused-resource cleanup — with estimated monthly savings per action. Cost attribution is mapped back to teams and projects.

## Key decisions

- **Deterministic answers first, LLM only where reasoning is required** — Most of what the report needs is arithmetic: day-over-day deltas, threshold breaches, cost per building, which query burned the most CPU. Handing those to a model buys nothing and introduces a way to be quietly wrong. The report writer computes them straight from the collected JSON. The LLM gets only the open-ended questions — optimization suggestions, risk assessment, the executive summary. Where both produce an answer, the deterministic one wins.
- **Three-phase pipeline, communicating through files on disk** — Collection, cross-reference, and persist as separate phases — collection nightly, deep analysis on the same run, recommendations in the merged report. Each phase writes JSON into a per-run directory, and the next phase reads those files. Any collector can be run standalone against yesterday's data, a failed run can be inspected after the fact, and the parallel execution mode reuses the exact same CLIs without a separate code path.
- **A doer–reviewer loop instead of one LLM pass** — An analyst agent writes the report; a second agent re-reads the raw data files and checks the claims against them, returning approved, revision-needed, or filtered. Capped at two rounds. This was the cheapest thing I found that caught confident numeric claims the data did not support — which is the failure mode that would have killed trust in the report fastest.
- **Anomaly detection against 30-day patterns, not a static bill alert** — Static cost thresholds trigger false alarms when the business grows. Thirteen anomaly checks learn spending patterns — spikes above the 30-day average, a query dominating database cost, sustained CPU saturation — and flag deviations, catching real waste without crying wolf. Configured 4-tier breach thresholds sit alongside that for resources that do have a known ceiling.
- **Reused the platform's existing email path rather than sending directly** — The report is inserted as an attachment record and a message is published to the topic the platform's existing mail worker already listens on. That meant matching an older .NET binary serialization format on the wire, which is ugly. The alternative was a second delivery mechanism with its own credentials, retry behaviour, and failure modes to operate.
- **Two execution modes over one set of collectors** — The sequential orchestrator is the simple path. A multi-agent runner layers batched parallelism and a live progress dashboard on top, and dispatches each LLM question as its own agent call instead of a sequential subprocess. It calls the same collector, analysis, and alert CLIs, so it is a scheduling layer, not a fork.

## Evaluation

There is no formal evaluation harness, and that is the real gap. What exists instead is a correctness guard rather than a measurement: the reviewer agent validates the analyst's claims against the raw data files, and any question that can be computed exactly is computed exactly rather than generated. That bounds how wrong the report can be. It does not tell me how useful it is. The 30% cost reduction came from acting on unused-resource flags, query improvements, and rightsizing — validated by the bill moving, not by an eval set.

## What I'd change

- **Automated remediation for low-risk items** — Currently generates recommendations but does not act. Auto-deleting clearly unused dev resources or auto-scaling idle services would compound savings without human intervention.
- **Multi-cloud from the start** — Built specifically for Azure. The collection layer abstractions were not clean enough to easily add AWS or GCP. The cost model should have been cloud-agnostic from day one.
- **Thresholds should have been learned, not configured** — Breach detection reads per-resource thresholds from a seeded table, which means someone has to know the right number in advance and remember to update it. A rolling baseline computed from the resource's own history would have needed no seeding and would have adapted as the platform grew.
- **Anomaly checks needed a feedback path from day one** — Thirteen checks fire into a report and nothing captures whether a given alert was worth reading. Without that signal there is no principled way to tune the thresholds, so tuning stayed guesswork.
- **Per-run output directories made trend queries harder than they should be** — Writing each run's JSON into its own timestamped directory was right for debuggability and wrong for anything that spans runs. Comparisons that should have been a SQL query became file walks.
- **The two execution modes drifted** — Sharing the CLIs kept the logic in one place, but the sequential and parallel paths still ended up with different failure and retry behaviour. One mode with a parallelism flag would have been less to keep in sync.
