稻草人新闻RSS 聚合阅读

← 返回 💻 编程 & 软件工程

RADAR: Catch gray failures with anomaly detection

Databricks 6 天前 www.databricks.com

Some of the most damaging outages are the ones your monitoring never flags: a slice of your customers quietly fails while every health check still reads normal. These "gray failures" leak users and revenue for hours before anyone connects the dots. This post is about catching them early with anomaly detection — how we do it at Databricks with a system called RADAR, and how you can build the same thing for whatever metric matters most to your business. It's written for the people who own service reliability: SREs, platform and data engineers, on-call responders, and the engineering leaders they answer to.

When everything is green but nothing is fine

Picture a normal Wednesday. Keeping a customer-facing service reliable is your job, and every dashboard on your wall is green — CPU healthy, latency fine, servers up, database connected. By every signal your team watches, the system looks perfect.

It isn’t.

  • 9:30 — A routine deploy slips a subtle bug into your checkout flow.
  • 9:35 — One in twenty customers paying by credit card silently fails. After a couple of retries, they give up and leave.
  • 12:40 — The first support ticket lands. It looks like just another mistyped card number, so nobody blinks.
  • 14:20 — Two more tickets arrive on the same issue.
  • 14:25 — Your support lead spots the pattern and escalates.
  • 16:00 — Engineers track down the bug and ship a fix.

For nearly seven hours, your monitoring insisted everything was fine while customers walked and revenue leaked.

What is a gray failure?

That Wednesday is a textbook gray failure. On the surface everything looks healthy; underneath, one specific piece has quietly stopped working — and it hurts customers without ever tripping an alert.

Two things make gray failures so sneaky:

  • They’re partial. It’s not everyone — just one slice, like a single type of credit card. There’s no server crash you’d catch instantly, just a messy middle where most users are fine and one group fails the whole time.
  • They grow. What starts as a handful of affected customers spreads. Left alone, more and more people hit the same wall.

Think of it as smoke behind the wall. From the outside the house looks fine, but inside the damage is spreading — and the longer you wait, the bigger the blast radius. Researchers have a name for the underlying problem, too: Microsoft’s Gray Failure: The Achilles’ Heel of Cloud-Scale Systems calls it differential observability — your failure detectors don’t notice a problem even while your users clearly do.

Why waiting for customer reports fails

Most teams handle gray failures exactly the way that Wednesday played out: they wait for customers to tell them. Customer reports matter — they’re real human pain — but your customers shouldn’t be your monitoring system. Leaning on reports alone has three problems:

  • It’s manual. Someone has to notice the same complaint across a pile of tickets. That’s easy to miss.
  • It’s delayed. By the time enough people complain for anyone to connect the dots, hours or days have passed.
  • It’s silent. Most affected customers never file a ticket at all. They just leave.

The fix isn’t to stop reading tickets — keep doing that. It’s to add automatic detection that runs all the time and catches what people miss. Concretely, you want something that fires the moment a lot more customers than usual start hitting the same issue at the same time.

Customer reports aloneAdd automatic detection
Manual — easy to missCatches what people miss
Delayed — noticed in daysFast — noticed in real time
Customers suffer in silenceFlags the spike — many users at once

在原文站打开 ↗

Cloudflare Workers 每 3 分钟抓一批,9 批轮完最快约 27 分钟 · 点右上 ↻ 立刻全量抓一次