Skip to Content

Your Data Isn't Good or Bad.

It is fit or unfit for a question.
October 1, 2026 by
Luis Roberto Aguirre Salazar

Before an AI pilot is approved, somebody in the room usually asks whether the data is good enough. It sounds like a sensible question. It is also one that almost nobody can answer, because in that form there is no answer.

Data is not good or bad in general. A customer file can be perfectly adequate for building a mailing list and useless for working out which accounts are likely to pay late. The same table passes one question and fails the next. Quality is not a property of the data. It is a relationship between the data and a specific question.

That distinction matters more than it seems. In most mid-sized companies, the data question is treated as a general clean-up: an initiative to tidy, deduplicate and standardise, launched ahead of any AI work, with no end date and no use case attached. It feels prudent. In practice it tends to produce tidier data that is still unfit for the question the business eventually asks, because nobody wrote that question down first.

My view is that the order should be reversed. Start with the question, then test whether the data can carry it.

Start with a decision, not a dataset

Underneath, an AI use case is a repeated decision: which invoice to chase first, which order to flag, where a request should be routed. The value comes from making that decision faster, more consistently or at greater volume. So, the first thing to write down is the decision itself, in a sentence a non-technical director can read: what is being decided, by whom, and from which information.

That sentence does two jobs. It names the fields the use case needs, which is usually a short list rather than “all our data”. And it gives the people who know the business something concrete to challenge. A decision written on paper exposes assumptions that a data inventory never will.

A short, hypothetical illustration: suppose the decision is which overdue invoices to contact first. The sentence might read: the collections lead decides each morning which accounts to call, using invoice age, amount, payment history and any open dispute. Four fields, one person, one moment. Now the data conversation is concrete. Is the dispute flag filled in by the time the list is built? Is payment history recorded per invoice, or only per customer? Those questions can be answered. “Is our data ready for AI?” cannot.

Three tests before the first use case

Once the decision is written down, three tests are worth running before anything is built. None of them requires technical knowledge. They require access to the people who know how the data is captured.

Timing. Is the information available at the moment the decision is made? Plenty of fields are complete in the database and still arrive too late: an outcome recorded weeks after the fact, a status updated in bulk at month end, a reason code filled in retroactively. A system cannot use at decision time what the business only learns afterwards.

Coverage. Does the data include the cases that matter, not only the easy ones? Records tend to be thorough for the routine path and thin for the exceptions: disputed invoices, returns, manual overrides, terms agreed by phone. If a use case is meant to handle exceptions, but exceptions are the records that were never properly captured, the data is silent exactly where the system will be asked to speak.

Consistency over time. Was the data recorded the same way across the period it will be used? Categories get renamed, fields get repurposed, a team starts using a free-text box for something the form was never designed for. Each change is reasonable when it happens. Together they mean a single column can carry several meanings, depending on the year and on who filled it in.

A check that costs almost nothing

Before any tooling, take the written decision and a small sample of real, recent records. Then ask the person who makes that decision today to make it using only the fields the use case will have. Wherever they hesitate, go looking elsewhere, or say that it depends on something that is not in there, the data is unfit for that question, at least as it is captured today.

This check is informal and does not replace a proper assessment. But in my view, it surfaces many of the problems that would otherwise appear only after the build, when they are far more expensive to deal with.

A failed test is a finding, not a failure

When the data does not pass, the instinct is often to postpone AI until the data has been fixed. That is rarely the best option. A failed test usually points to one of three paths: narrow the question to one the data can support today; change how the information is captured, so that the next months of records are usable; or choose a different first use case whose data already holds up.

Each of these is cheaper than discovering the gap midway through a build, after the decision to go ahead has already been made in public. It also gives the wider clean-up a shape. Instead of tidying everything, the company tidies what a specific decision needs and can tell when it is done.

The useful question is therefore not whether the data is good enough. It is: good enough for which decision, at which moment, and for which cases? A company that can answer that in writing is in a much better position to evaluate any tool it is offered, whatever it eventually chooses.

Rhisco’s Innovation & AI Lab builds AI agents and components to specification, with the appropriate governance architecture. If you are deciding where a first use case should start, contact us at contact@rhisco.com or visit rhisco.com/services.


This article was co-created with the assistance of artificial intelligence under strict supervision, editing, and verification of our team.

Specification-Driven Development
The record before the code