Composite health score, not metric-level alerts
Individual metrics in isolation cause alert fatigue. A composite 0–100 score per service synthesizes response time, error rate, throughput, and utilization into one number that is actionable.
Health scoring & regression detection
Catching service degradation before users notice it — nightly analysis of Azure Monitor metrics and SQL telemetry.
Stack
Built at Apra Labs
Problem
Solution
Outcome
0–100
health score per service
30-day
rolling baselines
Pre-impact
regression detection
Auto
root-cause analysis
System design · Pipeline · Decisions
System Architecture
How It Works
Telemetry
Azure Monitor data and SQL dynamic management views feed the scoring engine every night. Database telemetry is sampled at 15-second intervals from sys.dm_db_resource_stats, sys.dm_db_partition_stats, and Query Store. Four dimensions — response time, error rate, throughput, resource utilization — are computed per service, alongside seven-day peak metrics and Advisor recommendations.
Scoring
A composite health score (0–100) is calculated per service using weighted dimensions. Scores are compared against 30-day rolling baselines with outlier exclusion, so a service that is always slow on Mondays does not fire false alarms every Monday. Query cost attribution splits database CPU across the top queries, making a single runaway statement visible in the score rather than buried in a total.
Detection
Statistical deviation beyond configurable 4-tier thresholds triggers regression alerts. Thirteen anomaly checks run on the same pass: sustained CPU saturation, a table gaining 500MB overnight, a single query dominating database cost, cost and utilization spikes above the 30-day average. Root-cause analysis pinpoints which metric dimension drove the degradation. Breach results persist to SQL with 365-day retention, and the merged report goes out by email through the existing platform delivery path.
Key Decisions
Individual metrics in isolation cause alert fatigue. A composite 0–100 score per service synthesizes response time, error rate, throughput, and utilization into one number that is actionable.
30-day rolling baselines with outlier exclusion adapt to the service's natural patterns. A service that's always slow on Mondays doesn't fire false alarms every Monday. Configured 4-tier breach thresholds sit alongside that for resources that do have a known ceiling.
Azure Monitor tells you a database is hot. sys.dm_db_resource_stats and Query Store tell you which query made it hot, sampled every 15 seconds. Stitching those in the same nightly run is what turns "CPU is high" into "this schema change three weeks ago is eating a third of the bill."
The 13 anomaly checks and 4-tier breaches are arithmetic over collected JSON. Handing those to a model buys nothing and introduces a way to be quietly wrong. The LLM is reserved for the reasoning layer of the report — what to do about a regression, not whether one occurred.
Evaluation
There is no formal evaluation harness. What exists instead is a correctness guard: anomaly checks and health scores are computed from collected telemetry rather than generated, and a reviewer pass checks any LLM write-up against the raw files. That bounds how wrong a finding can be. It does not tell me how often a flagged regression was the one operators actually cared about — there is no feedback path from "was this alert worth reading."
What I'd Change
Get in Touch