Data Debt

Why Bad Data Beats Good Process

· Operations,Strategy

Someone says "let me get the real numbers before we go in" and the room moves on. Nobody asks what was wrong with the numbers already on the dashboard. The moment is treated as preparation, not a symptom.

That's the pattern Data Debt produces: a standing practice of compensating for data nobody can fully trust, described in language that makes it sound like diligence.

The instinct, when this surfaces, is to reach for better reporting or a new system. Both can help at the margins; neither fixes the underlying problem if the data feeding them remains unowned and undocumented. A cleaner dashboard connected to the same conflicting sources presents the conflict more neatly, but doesn't resolve it.

What's happening in most growing businesses is this: systems were built at different times, for different reasons, by different people. Each one carries its own definition of "customer," its own logic for what counts as revenue, and its own interpretation of headcount. Nobody ever merged those definitions because nobody was ever asked to. The systems worked, individually. The problem arrived when decisions needed to draw from all of them at once, and the definitions turned out not to match.

Good process built on data like that makes things worse. A well-run decision-making rhythm, applied to numbers that diverge at their sources, moves the wrong answer through the business faster and with more confidence attached to it than a messy process would have generated. A mediocre process on good data will produce better decisions than a good process on bad data.

That's Data Debt: the accumulated cost of data nobody can trust, and the compounding effect of decisions made against it.

That makes it a leadership problem. Data quality tends to get handed to whoever looks after the systems, but the systems team isn't the one deciding on pricing, hiring or investment on the strength of a number. Leaders are, and each of those decisions is only as sound as the figure behind it.

Recognising Data Debt in Your Business

The clearest markers are found in the things people say.

"That depends which system you're looking at" is one of them, offered as a normal answer to a normal question, without a flicker of concern. Two people in the same meeting quote different figures for the same metric; neither finds that surprising. "Let me get you the real numbers" has become a standing ritual before anything consequential gets decided, and the phrase lands without irony because everyone in the room understands what it means.

The process markers follow a similar pattern. A recurring decision depends on someone manually reconciling a spreadsheet, and their absence would stall it. One person, often without this being any formal part of their job, has become the de facto authority on which system is actually right; everyone defers to them rather than to anything documented. Reports get cross-checked against a second system out of habit, because the first isn't fully trusted, and nobody has ever asked why that habit started.

The ownership markers explain why the others persist. Nobody owns the question "what does this field actually mean" as a standing responsibility. Definitions live in individual heads, built up through experience and ongoing, informal correction. New starters learn "how we actually read the numbers here" from a colleague in passing rather than from anything written down. When that colleague leaves, so does the knowledge. What looks like institutional memory is often one person's undocumented interpretation, which becomes the default until someone else's interpretation takes over.

Mantage's framework primer for this series, "Organisational Debt: Primer"[1], offers a diagnostic that applies here. Take the three most critical questions your business answers regularly: the ones that shape how you hire, what you sell, where you invest. Ask whether your team can answer each one confidently, from current systems, in under five minutes, with a clear source they can name. A confident answer with a named source is a good sign. A pause, a caveat, a trip to a spreadsheet, or "let me check with Sarah" is the clearest marker of Data Debt in practice.

None of this requires a data background to spot. It requires noticing what's already routine and asking, for the first time, why it's routine.

The Cost Curve of Data Debt

Data Debt compounds in a way that makes it difficult to see from inside the business.

Each new integration adds another place a definition can diverge. A second CRM brought in to serve a different team; a finance tool adopted for reporting that draws from a different source than the one sales uses; a headcount figure that lives in HR software, in a spreadsheet, and in someone's quarterly deck, and hasn't matched across all three for longer than anyone can remember. As the business grows, the number of places where a single figure can become two figures multiplies.

Manual reconciliation, adopted as a stopgap, rarely gets revisited once it starts working. What begins as a temporary fix to bridge two systems becomes permanent, invisible labour. Someone is spending time each week doing work that only exists because the data underneath it was never properly governed. None of that time appears in a budget or a project plan, so nobody ever weighs it against anything else. Each new system the business brings in adds one more source for them to reconcile by hand.

Confidence and accuracy drift apart as the business scales. Reports get produced faster; formats get more polished; presentations get more assured. Checking whether those figures are right gets harder with every new source, and happens less often. A number in a board pack carries authority because of how it's presented: the table, the chart, the percentage change formatted to one decimal place. The question of whether the same figure, run by a different person on a different system, would match doesn't get asked.

AI doesn't correct for any of this. A model connected to undocumented, unreconciled data inherits whatever quality already exists and scales it, faster and with more apparent authority than any person reconciling a spreadsheet ever had. It produces confident, fluent output. That confidence is the risk: there's no visible hesitation, no caveat, no sense that the answer is only as reliable as the inputs behind it. The wrong answer arrives looking exactly like the right one.

The good news is that fixing Data Debt doesn't require fixing all data everywhere. It requires identifying the few numbers that drive recurring decisions and governing them deliberately, rather than leaving governance to accumulate by habit.

Reducing Data Debt: Where to Start

Tool 1: Name a Single Source of Truth for Each Critical Number

For each metric that recurring decisions depend on, name which system's figure is authoritative when two disagree, and record that decision somewhere everyone can find it. This doesn't require one system for everything, just a decision made once and written down per number that matters. Without a named source of truth, every disagreement gets re-litigated from scratch, in every meeting, by whoever happens to be in the room.

Tool 2: Write Down What the Fields That Matter Actually Mean

Build a short working definition for each critical metric from Tool 1: what counts, what doesn't, where it's measured from. A page is enough; a business-wide data dictionary rarely gets finished. A number without a documented definition means whatever the reader that day assumes it means, and those assumptions drift apart over time, unnoticed, until someone asks the question in public and everyone gives a different answer.

Tool 3: Treat Manual Reconciliation as a Flag, Not a Workflow

Wherever someone is reconciling numbers by hand before a decision can be made, treat that as a symptom to fix, and resist the urge to make the reconciliation itself faster. Give it an owner and a deadline to resolve the mismatch properly, or to confirm that manual reconciliation is the right long-term answer. Done well, it feels like competence. It's actually the cost of Data Debt, paid week after week by whoever's doing it.

Tool 4: Check Trust Before You Scale Anything on Top of a Number

Before automating a process or building an AI step on a metric, confirm the team already trusts that number without a manual check. If the answer is "not quite," fix the data first. Automating an untrusted number delivers the same wrong answer faster. Speed applied to untrustworthy data multiplies the problem, and AI adds more speed than anything else.

The Realities of Addressing Data Debt

Addressing Data Debt is an ongoing habit. New systems, new hires, and new definitions keep arriving, and every new integration brings its own field names and assumptions about what counts. The discipline has to keep pace with the business, so it belongs in how decisions get made, with no end date.

For most small and medium enterprises, full data governance (a formal data warehouse, a dedicated data team, an enterprise governance function) is neither realistic nor necessary. The aim is a small, named set of trusted numbers: the ones decisions depend on. Everything else can remain imperfect without causing meaningful harm. The discipline is knowing which numbers can't.

Fixing the data sometimes means admitting it's been wrong. Naming Data Debt can surface a figure that's been quoted confidently and incorrectly for months, or longer. A number that's been wrong in every board pack for two years doesn't become right by ignoring it. The discomfort of naming the error is a one-off; the cost of not naming it keeps compounding.

Leaders can't personally verify data quality any more than they can audit an engineering codebase or check a compliance claim from a specialist they rely on. That's fine. What they can do is ask three questions of any number that matters: who owns it, how its definition was agreed, and when anyone last checked that it still holds. None of those needs a data background to ask, and the answers reveal more than whether the figure looks plausible.

Data Debt in the Wider Organisation

Data Debt is easily confused with Technical Debt[6]. Technical Debt is about engineering systems and codebases: the structural choices made under pressure that make future changes harder and more expensive. Data Debt is about what those systems hold and whether it can be trusted, independent of how well-engineered the underlying infrastructure is. A company can run clean, well-maintained software on top of fifteen years of inconsistently defined fields, or keep a legacy system held together with workarounds that somehow produces a number everyone trusts. The engineering quality and the data quality move independently.

Operational Debt[3] is the other near neighbour. Operational Debt is undisciplined business process and tooling in general: the workarounds, the ad-hoc tools, the procedures that were improvised and never formalised. Data Debt is specifically about whether the numbers those processes produce can be trusted. A business can tighten its operations and still carry Data Debt if nobody has decided which figure is authoritative or what the fields mean.

That distinction matters particularly when choosing how to handle a process step. The piece "Deterministic by Default"[8] argues for being deliberate about when a rule-based method is the right choice and when an AI approach earns its place. But that choice doesn't resolve much if the inputs to whichever method is chosen are already compromised. Trustworthy data is a precondition for that kind of discipline: the right tool pointed at the wrong number still produces the wrong answer.

In Mantage's operations review, mapping the data landscape runs alongside reviewing process and governance, rather than treating "what do we actually trust" as a separate specialist exercise. The question of which numbers the business runs on belongs in the same conversation as how it runs. In strategy delivery, the numbers behind a plan are checked before anything is built on them: a strategy built on a figure nobody can defend is a strategy built on a guess presented as evidence. Mentoring supports the harder conversation Data Debt tends to avoid, naming out loud that a long-quoted number might be wrong, without it turning into blame.

This is the eighth piece in the Organisational Debt series, following the primer[1] and the articles on Cultural Debt[2], Operational Debt[3], Capability Debt[4], Strategy Debt[5], Technical Debt[6], and Regulatory Debt[7]. Innovation Debt is the only type from the original eight left to publish after this.

Data doesn't have to be perfect to be trusted. It has to be owned, defined, and checked on a rhythm, for the handful of numbers the business runs on. Every process, plan, or AI project built on top of those numbers is only as reliable as they are.

Referenced Articles

Ady Coles helps organisations reduce operational friction so strategy has a chance to work. He focuses on operational clarity, sensible governance, and the thoughtful use of automation; not optimisation for its own sake, but making work easier, decisions clearer, and scale more sustainable as organisations grow.