Back

Infrastructure Performance Agent

Health scoring & regression detection

Catching service degradation before users notice it — nightly analysis of Azure Monitor metrics and SQL telemetry.

Stack

TypeScriptNode 22 / tsxAzure MonitorAzure SQL DMVsQuery StoreAzure Service BusStatistics

Built at Apra Labs

Problem

  • Users reporting issues before ops team aware
  • Slow performance regressions going undetected
  • A schema change can make one query eat a third of database CPU
  • Metric-based alerts causing severe alert fatigue
  • No composite health visibility across services
  • Root cause analysis done manually each time — SLA breaches discovered after the fact

Solution

  • Composite 0–100 health scoring per service across four dimensions
  • 30-day rolling baseline anomaly detection with outlier exclusion
  • Database telemetry sampled at 15-second intervals from SQL DMVs and Query Store
  • 13 anomaly checks: CPU saturation, table growth, query cost dominance, peak metrics
  • 4-tier breach thresholds plus automated root-cause analysis
  • Nightly report through the existing platform mail path, with live progress in parallel mode

Outcome

  • Pre-impact regression detection running live
  • Per-service health scores visible to all teams
  • Alert fatigue significantly reduced across ops
  • Degradation caught well before user impact
  • MTTR reduced with automated root-cause hints
  • SLA compliance tracking now fully automated

0–100

health score per service

30-day

rolling baselines

Pre-impact

regression detection

Auto

root-cause analysis

Architecture Review

System design · Pipeline · Decisions

System Architecture

How It Works

  1. 1

    Telemetry

    Azure Monitor data and SQL dynamic management views feed the scoring engine every night. Database telemetry is sampled at 15-second intervals from sys.dm_db_resource_stats, sys.dm_db_partition_stats, and Query Store. Four dimensions — response time, error rate, throughput, resource utilization — are computed per service, alongside seven-day peak metrics and Advisor recommendations.

  2. 2

    Scoring

    A composite health score (0–100) is calculated per service using weighted dimensions. Scores are compared against 30-day rolling baselines with outlier exclusion, so a service that is always slow on Mondays does not fire false alarms every Monday. Query cost attribution splits database CPU across the top queries, making a single runaway statement visible in the score rather than buried in a total.

  3. 3

    Detection

    Statistical deviation beyond configurable 4-tier thresholds triggers regression alerts. Thirteen anomaly checks run on the same pass: sustained CPU saturation, a table gaining 500MB overnight, a single query dominating database cost, cost and utilization spikes above the 30-day average. Root-cause analysis pinpoints which metric dimension drove the degradation. Breach results persist to SQL with 365-day retention, and the merged report goes out by email through the existing platform delivery path.

Key Decisions

Composite health score, not metric-level alerts

Individual metrics in isolation cause alert fatigue. A composite 0–100 score per service synthesizes response time, error rate, throughput, and utilization into one number that is actionable.

Rolling baselines over fixed thresholds

30-day rolling baselines with outlier exclusion adapt to the service's natural patterns. A service that's always slow on Mondays doesn't fire false alarms every Monday. Configured 4-tier breach thresholds sit alongside that for resources that do have a known ceiling.

SQL telemetry in the same pipeline as Monitor metrics

Azure Monitor tells you a database is hot. sys.dm_db_resource_stats and Query Store tell you which query made it hot, sampled every 15 seconds. Stitching those in the same nightly run is what turns "CPU is high" into "this schema change three weeks ago is eating a third of the bill."

Deterministic checks first, LLM only for the write-up

The 13 anomaly checks and 4-tier breaches are arithmetic over collected JSON. Handing those to a model buys nothing and introduces a way to be quietly wrong. The LLM is reserved for the reasoning layer of the report — what to do about a regression, not whether one occurred.

Infrastructure Performance Agent — Architecture

Infrastructure Performance Agent architecture

Evaluation

There is no formal evaluation harness. What exists instead is a correctness guard: anomaly checks and health scores are computed from collected telemetry rather than generated, and a reviewer pass checks any LLM write-up against the raw files. That bounds how wrong a finding can be. It does not tell me how often a flagged regression was the one operators actually cared about — there is no feedback path from "was this alert worth reading."

What I'd Change

  • Cross-service correlation
  • Predictive scoring
  • Anomaly checks needed a feedback path from day one
  • Breach thresholds should have been learned, not seeded

Get in Touch

Let's build
something great.