The Real Cost of a Wrong Metric
For a stretch of my time at Meta, my job was to make sure no incorrect external metric was ever published on any Meta platform — Facebook, Instagram, Messenger, Ads, Video, and close to twenty other product orgs. These weren't obscure internal numbers; they were the ones paid advertisers actually looked at — link clicks, video views, impressions — the core signals businesses relied on to judge whether their ad spend, and their business, was actually working. Said out loud, that sounds like a QA function. It isn't. A wrong external metric at that scale isn't a bug ticket; it's a restatement, a headline, a board conversation — sometimes a multi-million-dollar problem before anyone even agrees on how it happened.
None of it held up on manual review alone. At that scale, it took building the right processes, automation, and tooling — including AI — to catch drift and anomalies across that many metrics and pipelines before a person ever had to look. And it wasn't always the metric itself that was wrong; just as often, the real issue lived one layer down, in the infrastructure the metric depended on — the pipeline, the job, the system feeding it — not the definition sitting on top of it. Tooling could catch that kind of break. It couldn't catch the more common one.
Why Metrics Go Wrong
The uncomfortable lesson: most metric errors aren't caused by bad code. They're caused by ambiguity nobody resolved before the metric shipped. Two teams define "active user" slightly differently. A pipeline change quietly shifts a join, and the number moves half a percent — small enough that nobody questions it, which is exactly why it does the most damage: it's believable enough to act on.
A Recurring Review Cycle
Fixing that isn't a one-time audit. It's a recurring cycle: going back through the metrics that mattered most and reviewing them again on a set cadence, metric by metric, to catch inconsistency before it has a chance to compound.
That started with confirming who actually owned each metric, since ownership drifts as teams reorg and people move on, and a metric with no clear owner is a metric nobody's watching. From there, it meant coordinating across every team that touched it and working through a checklist of nearly twenty things per metric, including:
- The business definition — what the metric is actually supposed to mean
- The technical definition — how it's actually computed
- The data pipeline behind it
- The SLA it was supposed to meet
- Consistency across platforms — the same "active user" should mean the same thing on Facebook and on Instagram
...and a dozen more checks like it, each one a place a number could quietly drift without anyone noticing.
Getting Buy-In From Busy Teams
The part that actually took the most convincing wasn't getting agreement in principle — most people nodded along that metric definitions should be a contract, not a suggestion. It was getting engineers and product teams, already stretched across their own roadmaps, to actually spend cycles on it. The lever that worked was data, not persuasion: bringing the actual cost and risk numbers to the table so metric-review work could compete for priority on the same terms as everything else on the roadmap, instead of as a vague "important" ask.
Winning that argument took two different kinds of conversations, depending on the room: program health updates to VPs and Directors, and separately, pitches to specific business groups on why the review work belonged on their roadmap at all. Different audience, same throughline — lead with the cost and risk numbers, not the ask.
The Actual Job
When people stop trusting a number, they slow down. They double-check everything before deciding, so things just take longer. But that's not even the worst part — there's the opportunity cost of decisions that get delayed or never happen at all. And customers start trusting you less when the numbers turn out to be wrong. None of that shows up in a bug report.
The teams that get this right don't have fewer bugs than everyone else. They just have faster, more boring incident response — because they built the observability and the audit discipline before they needed it, not after a metric embarrassed someone in a leadership review.
If you're the one trying to convince your company to invest in this, lead with the organizational fix, not the technical one — pipelines and dashboards are the easy part. Getting a company to agree on what a number means, and to keep agreeing as teams and owners change, is the actual job. Go back and review the infrastructure and the metrics on a regular cadence — not just when something breaks. And that review doesn't have to be entirely manual anymore — AI and agents can now handle a real share of that checking, catching drift before a person has to go looking for it. Code doesn't need convincing. People do.
Data Engineering & Analytics Leader with 16+ years building and scaling data platforms, analytics ecosystems, and AI-driven solutions. More about my background →