Back

Cloud Cost Intelligence Agent

Automated FinOps analysis

Finding out why the cloud bill moved — every night, before anyone has to ask. Identified 30% savings across the Azure estate.

Stack

TypeScriptNode 22 / tsxAzure Cost ManagementAzure Resource GraphAzure AdvisorAzure MonitorAzure SQLClaude SonnetAzure Service Bus

Built at Apra Labs

What made it hard

Cloud spend was growing month over month and the billing portal could only say that the total had moved, never why. Answering "why" meant stitching cost data, telemetry and resource inventory together by hand, so it happened quarterly at best, far too slow while waste compounded silently. A schema change could hand a third of database CPU to one query and nobody would notice for a quarter.

Problem

  • Cloud spend growing unchecked month over month
  • The billing portal shows the total went up, not why
  • A schema change can make one query eat a third of database CPU
  • Detached disks and idle resources still being paid for months later
  • Answering "why" meant stitching cost, telemetry, and inventory by hand
  • Quarterly manual audits far too slow — waste compounding silently

Solution

  • 19 collectors across 3 dependency phases — six Azure APIs, SQL DMVs, and the app database
  • 13 deterministic anomaly checks and 4-tier breach thresholds
  • 20+ standing questions answered from collected data; LLM only for reasoning
  • Doer–reviewer loop: an analyst writes, a second agent checks claims against raw files
  • Query cost attribution splits the database bill across top queries by CPU share
  • Rightsizing, reservation, and unused-resource recommendations, emailed through the existing platform mail path

Outcome

  • 30% cloud cost reduction achieved in months
  • Unused resources automatically flagged for cleanup
  • Prioritized optimization recommendations delivered nightly
  • Teams now own and track their spend
  • Anomaly detection over 15-second telemetry samples, reported every night
  • Millions saved across the organization annually

30%

cost reduction

19

data collectors

13

anomaly detectors

3

analysis phases

Architecture Review

System design · Pipeline · Decisions

System Architecture

How It Works

  1. 1

    Collect

    Twelve independent collectors pull cost breakdowns, 30-day trends, resource inventory, orphaned disks and NICs, reservation utilization, Advisor recommendations, seven-day peak metrics, and database telemetry sampled at 15-second intervals. No cross-dependencies, so they parallelize freely. Data is normalized into a common cost model and written as JSON into a per-run directory.

  2. 2

    Cross-reference

    Five more collectors read that output back off disk and join it. Query cost attribution splits the database bill across the top queries by CPU share. Breach detection compares each resource against 4-tier configured thresholds. The anomaly detector runs 13 checks: cost spikes above the 30-day average, sustained CPU saturation, a table gaining 500MB overnight, a single query dominating database cost, a reservation about to expire.

  3. 3

    Persist

    Breach results and telemetry history land in eight SQL tables. A retention pass trims everything past 365 days on each run.

  4. 4

    Analyze

    The deterministic report writer and the LLM loop run in parallel. Arithmetic — day-over-day deltas, threshold breaches, cost per building, which query burned the most CPU — is computed straight from the collected JSON. The LLM gets only the open-ended questions: optimization suggestions, risk assessment, the executive summary. A reviewer agent re-reads the raw files and checks the claims, capped at two rounds. Where both produce an answer, the deterministic one wins.

  5. 5

    Deliver

    The merged report goes out by email through the existing platform delivery path — an attachment record and a Service Bus message the platform's mail worker already listens on. Recommendations are prioritized: right-sizing, scheduling, reservation purchases, unused-resource cleanup — with estimated monthly savings per action. Cost attribution is mapped back to teams and projects.

Constraints

  • Attribution has to be defensible: a wrong cost claim sends an engineer down a multi-day dead end
  • Six Azure APIs, SQL DMVs and the app database all had to agree before a number could be reported
  • The analysis itself runs nightly and cannot cost a meaningful fraction of what it saves
  • An LLM cannot be trusted to compute the numbers, only to reason about them once computed
  • Findings land in engineers inboxes via the existing platform mail path — no new system to adopt

Key Decisions

Deterministic answers first, LLM only where reasoning is required

Most of what the report needs is arithmetic: day-over-day deltas, threshold breaches, cost per building, which query burned the most CPU. Handing those to a model buys nothing and introduces a way to be quietly wrong. The report writer computes them straight from the collected JSON. The LLM gets only the open-ended questions — optimization suggestions, risk assessment, the executive summary. Where both produce an answer, the deterministic one wins.

Three-phase pipeline, communicating through files on disk

Collection, cross-reference, and persist as separate phases — collection nightly, deep analysis on the same run, recommendations in the merged report. Each phase writes JSON into a per-run directory, and the next phase reads those files. Any collector can be run standalone against yesterday's data, a failed run can be inspected after the fact, and the parallel execution mode reuses the exact same CLIs without a separate code path.

A doer–reviewer loop instead of one LLM pass

An analyst agent writes the report; a second agent re-reads the raw data files and checks the claims against them, returning approved, revision-needed, or filtered. Capped at two rounds. This was the cheapest thing I found that caught confident numeric claims the data did not support — which is the failure mode that would have killed trust in the report fastest.

Anomaly detection against 30-day patterns, not a static bill alert

Static cost thresholds trigger false alarms when the business grows. Thirteen anomaly checks learn spending patterns — spikes above the 30-day average, a query dominating database cost, sustained CPU saturation — and flag deviations, catching real waste without crying wolf. Configured 4-tier breach thresholds sit alongside that for resources that do have a known ceiling.

Reused the platform's existing email path rather than sending directly

The report is inserted as an attachment record and a message is published to the topic the platform's existing mail worker already listens on. That meant matching an older .NET binary serialization format on the wire, which is ugly. The alternative was a second delivery mechanism with its own credentials, retry behaviour, and failure modes to operate.

Two execution modes over one set of collectors

The sequential orchestrator is the simple path. A multi-agent runner layers batched parallelism and a live progress dashboard on top, and dispatches each LLM question as its own agent call instead of a sequential subprocess. It calls the same collector, analysis, and alert CLIs, so it is a scheduling layer, not a fork.

Cloud Cost Intelligence Agent — Architecture

Cloud Cost Intelligence Agent architecture

Evaluation

There is no formal evaluation harness, and that is the real gap. What exists instead is a correctness guard rather than a measurement: the reviewer agent validates the analyst's claims against the raw data files, and any question that can be computed exactly is computed exactly rather than generated. That bounds how wrong the report can be. It does not tell me how useful it is. The 30% cost reduction came from acting on unused-resource flags, query improvements, and rightsizing — validated by the bill moving, not by an eval set.

What I'd Change

  • Automated remediation for low-risk items
  • Multi-cloud from the start
  • Thresholds should have been learned, not configured
  • Anomaly checks needed a feedback path from day one
  • Per-run output directories made trend queries harder than they should be
  • The two execution modes drifted

Get in Touch

Let's build
something great.