Data quality consulting: what small firms actually need
A quote line matched Clematis Armandii Apple Blossom against Camellia Japonica, because the two cultivar names looked similar on the page. Completely the wrong plant, at completely the wrong price.
The mistake was in the data before any software went near it. It went through because nothing in the process was looking for it.
That is what data quality consulting is really about, and it is a long way from what the big consultancy pages are selling. They are written for companies with a data team. This is the version for a firm where the person who knows the records also answers the phone.
What data quality actually means
Data quality is whether the records you run on can be trusted for the job in hand. Not perfect. Trusted for that job.
A customer list good enough to post a Christmas card to is not good enough to run a direct debit from. The same list, the same day, two different answers.
In practice it comes down to a handful of plain questions. Does each customer appear once. Is a date always a date. Do the prices in the system match the invoices you actually sent. Are the fields you depend on filled in, or empty on a third of the records.
You do not need a maturity model to answer those. You need an hour and somebody who knows what the fields are supposed to mean.
The number every guide quotes, and why this one does not
Every page ranking for this term quotes the same Gartner estimate of what poor data quality costs the average business. It is usually repeated without a date, and it has been passed between vendor blogs for years.
I am not going to repeat it. Not because it is wrong, but because I cannot check it, and the rule here is that every figure names a source I have read.
Here is one I can stand behind. Microsoft's 2026 Work Trend Index, a survey of 20,000 knowledge workers, found 86 percent of AI users treat what it produces as a starting point rather than a final answer. Half of them named quality control of AI output as a skill that is becoming more important.
Microsoft sells the tools being measured, so read it with that in mind.
The useful part is not the percentage. It is that checking has quietly become the work.
What AI does to bad data
It makes it faster and more confident.
The plant line above is the clearest example I have. A human typed a cultivar name that looked close enough. The matching stage accepted it. Nothing downstream asked whether a climber and a shrub were really the same thing, so a wrong plant at a wrong price travelled all the way to a quote.
AI did not invent that error. It carried it, quickly and without hesitating.
This is the bit worth being blunt about, and it is the same argument I make about AI in an accounts practice. Feeding a language model a spreadsheet full of duplicates does not clean the spreadsheet. It produces a confident summary of duplicated figures, in a tone that sounds like it has checked.
Which is why the scarce skill is no longer producing the work. It is judging whether the work is right, and that judgement sits with the person who knows that agapanthus and African lily are the same plant, and that a ninety centimetre tree fern is not a fifty centimetre one.
Where the errors actually come from
Almost never from one dramatic failure. They arrive quietly, from four or five ordinary places.
Two systems that do not talk, so somebody re-keys between them and drops a digit at five o'clock on a Friday. If that is your situation, accounting software integration fixes the cause rather than the symptom.
There is a compliance edge to it too. HMRC expects digital links between the pieces of software making up your VAT records, and retyping figures between systems is not a digital link. That is set out in VAT Notice 700/22.
Free text fields where a rule should be. If the address is a box anyone can type into, you will eventually have Altrincham, altrincham, ALT and Altricham.
The same customer created twice, because the second booking came in by phone and nobody searched first.
Fields that were optional when the system was set up and turned out to matter later. Nobody backfills them. They simply stay empty, and every report quietly excludes those rows.
And leavers. The person who knew that the codes beginning with a Z are historic has gone, and the knowledge went with them.
The checks worth doing first
Do these on one table, not on everything you own.
Count the duplicates. Sort by name, then by postcode, then by phone number, and look at how many near matches appear. This one number usually settles the argument about whether there is a problem.
Check the formats. Dates are the worst offender, because a system will happily hold two formats at once and neither looks wrong on screen until you sort by them.
Count the blanks in the fields you actually use. Not every field. The four or five that decide whether somebody gets invoiced.
Look for the dead records. Customers who have not bought since 2019, staff who left, prices for products you stopped selling. They are not harmful in themselves, but they make every count wrong.
Then pick the worst one and fix its cause. Not all of them. The worst one.
Who owns the customer record
This is what the big consultancy pages call governance, and it is the part small firms skip because the word sounds like it needs a committee.
At your size it is one question. Who is allowed to decide what a customer record should say, and who do the rest of you ask when it is wrong.
Write the answer down. One name per record type: customers, prices, staff, suppliers. That is the whole framework, and it does more good than a policy document nobody opens.
There is a legal edge to this as well. Under UK GDPR, personal data has to be accurate and kept up to date, which is the fourth data protection principle and not a matter of preference. The ICO's guidance on accuracy is short and plain, and worth ten minutes if you hold customer addresses.
Keeping it clean, which is the hard part
A clean up is a day's work. Staying clean is a habit, and habits are where this falls over.
The trick is to move the check to the moment the mistake is made. A validation rule that will not let a booking save without a postcode costs nothing and prevents the whole problem. Finding the same gap three weeks later costs somebody an afternoon.
Where a check cannot sit at the point of entry, put it in a report instead. A short weekly list of exceptions, sent to the person who owns that record type: records missing a required field, customers created twice this week, prices that do not match the last invoice.
That is the sort of thing automated business reporting is for, and a better use of it than another dashboard nobody opens.
The measure of success is boring. The list gets shorter, then stays short.
What this costs, and the tools question
Start with the tools, because most people ask about them first and they matter least.
You probably do not need deduplication software. The products in this market are built for organisations with millions of records, priced accordingly, and they will not tell you which of two addresses is the right one. A spreadsheet, a sort and an hour will find the duplicates in a list of four thousand customers.
What costs money is the fixing, and that is usually because two systems disagree.
Our one off work starts from £500, which is where a single clean up and a set of checks normally sits. Linking two systems so the problem stops coming back is from £500 for a single connection, rising for more involved builds. Ongoing checks delivered as a monthly report are from £75 a month.
Everything is fixed and agreed before it starts. If you want the longer version of how two systems end up disagreeing in the first place, digital transformation services for a small firm follows a business where the booking system and the accounts package had never reconciled.
What you could do on Monday
Open the table you would least like to be wrong. For most firms that is the customer list.
Sort it by postcode and read the first fifty rows. Not with a tool. With your eyes.
You are looking for three things: the same customer twice, a field that should never be empty and is, and a value that is obviously a typing mistake. Count what you find, and write the number down.
If it is under five, you do not have a data quality problem and you can stop reading. If it is over twenty, you have found your first job, and you now know something about your business you did not know at nine o'clock.
When not to hire us
If you have one system, and one person who knows it, and nobody is retyping anything into anything else, there is nothing here worth paying for. Do the Monday exercise, fix what you find, and keep your money.
The same applies if the answer is a setting rather than a project. A required field, a validation rule, a duplicate check in software you already own: those are twenty minutes with the manual, not an engagement.
And there is a limit to what any of this fixes. Cleaning a customer list does not tell you which of two addresses is current, and neither do we. Somebody who knows the customer has to decide.
What good checks do is find the twenty records that need deciding, instead of leaving them buried in four thousand that do not.
That is a smaller claim than the one on the big consultancy pages, and it is the one I can actually deliver.
If you are not sure whether your data is bad enough to bother with, ring on 0161 883 7818 and describe what keeps going wrong. If the honest answer is that you do not need us, that is what you will get.