Deterministic answers first, LLM only where reasoning is required
Most of what the report needs is arithmetic: day-over-day deltas, threshold breaches, cost per building, which query burned the most CPU. Handing those to a model buys nothing and introduces a way to be quietly wrong. The report writer computes them straight from the collected JSON. The LLM gets only the open-ended questions — optimization suggestions, risk assessment, the executive summary. Where both produce an answer, the deterministic one wins.
Three-phase pipeline, communicating through files on disk
Collection, cross-reference, and persist as separate phases — collection nightly, deep analysis on the same run, recommendations in the merged report. Each phase writes JSON into a per-run directory, and the next phase reads those files. Any collector can be run standalone against yesterday's data, a failed run can be inspected after the fact, and the parallel execution mode reuses the exact same CLIs without a separate code path.
A doer–reviewer loop instead of one LLM pass
An analyst agent writes the report; a second agent re-reads the raw data files and checks the claims against them, returning approved, revision-needed, or filtered. Capped at two rounds. This was the cheapest thing I found that caught confident numeric claims the data did not support — which is the failure mode that would have killed trust in the report fastest.
Anomaly detection against 30-day patterns, not a static bill alert
Static cost thresholds trigger false alarms when the business grows. Thirteen anomaly checks learn spending patterns — spikes above the 30-day average, a query dominating database cost, sustained CPU saturation — and flag deviations, catching real waste without crying wolf. Configured 4-tier breach thresholds sit alongside that for resources that do have a known ceiling.
Reused the platform's existing email path rather than sending directly
The report is inserted as an attachment record and a message is published to the topic the platform's existing mail worker already listens on. That meant matching an older .NET binary serialization format on the wire, which is ugly. The alternative was a second delivery mechanism with its own credentials, retry behaviour, and failure modes to operate.
Two execution modes over one set of collectors
The sequential orchestrator is the simple path. A multi-agent runner layers batched parallelism and a live progress dashboard on top, and dispatches each LLM question as its own agent call instead of a sequential subprocess. It calls the same collector, analysis, and alert CLIs, so it is a scheduling layer, not a fork.