RADAR: Catch gray failures with anomaly detection
AI Classified by Officially
Some of the most damaging outages are the ones your monitoring never flags: a slice of your customers quietly fails while every health check still reads normal. These "gray failures" leak users and revenue for hours before anyone connects the dots. This post is about catching them early with anomaly detection — how we do it at Databricks with a system called RADAR, and how you can build the same thing for whatever metric matters most to your business. It's written for the people who own service reliability: SREs, platform and data engineers, on-call responders, and the engineering leaders they answer to.
When everything is green but nothing is fine
Picture a normal Wednesday. Keeping a customer-facing service reliable is your job, and every dashboard on your wall is green — CPU healthy, latency fine, servers up, database connected. By every signal your team watches, the system looks perfect.
For nearly seven hours, your monitoring insisted everything was fine while customers walked and revenue leaked.
This is an extract. The publication continues at the source.
Read the original at the source: https://www.databricks.com/blog/radar-catch-gray-failures-anomaly-detection
Officially imported this from Databricks’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.
Provenance
- Organization
- Databricks — imported from official source
- Official source
- https://www.databricks.com/feed RSS
- Imported
- September 20, 2026 19:52
- Versions
- 1 recorded
- Identity
https://www.databricks.com/blog/radar-catch-gray-failures-anomaly-detection