Here is a pattern worth noticing. Ask why a procurement initiative failed and you will usually get a specific answer: the tool was wrong, the supplier underperformed, the team was too busy.
Ask a few more questions and you often reach the same place. Nobody could see the spend clearly. The supplier list had the same company three times. The contract terms were in an email folder.
Procurement data is not a technical subject. It is the thing that decides what your procurement function is capable of.
The four datasets
Almost everything procurement does rests on four sets of records. Most companies have all four in some form. Few have them in good condition.
| Dataset | What it holds | What it unlocks when it is good |
|---|---|---|
| Spend | What you bought, from whom, when, in what category | Consolidation, negotiation leverage, tariff exposure modelling, budget forecasting |
| Supplier master | One record per company: legal entity, bank details, status, contacts | Risk assessment, accurate total spend per supplier, fewer duplicate payments |
| Contracts | Prices, terms, end dates, notice periods, obligations | Renewals caught in time, invoice prices checked against agreed prices |
| Items and catalogue | What things actually are: descriptions, specifications, part numbers | Standardisation, catalogue buying, finding the same item bought twice at two prices |
The third column is the point. None of those outcomes is a data project. They are commercial outcomes that happen to be impossible without the data.
The evidence that data is the binding constraint
This is not just our opinion. It shows up in every recent study of why procurement initiatives stall.
Deloitte asked more than 250 chief procurement officers what they saw as the main internal risks in deploying generative AI. Data quality came top, named by 43.97%, ahead of governance, security and compliance concerns at 35.41%.
The 2026 ProcureCon CPO report found data quality and cross-system integration cited as a barrier by 54% of procurement leaders.
And the visibility gaps are measurable. Ardent Partners puts average spend under management at about 71%, meaning nearly three pounds in every ten are spent without procurement seeing them. McKinsey found 95% of companies can see risks in their direct suppliers but only 42% can see into the next tier.
The cost of poor data is not usually a dramatic failure. It is a tax on everything: every analysis takes longer, every negotiation starts from a weaker position, every risk question gets an approximate answer.
What good data buys you, concretely
Negotiating from facts
If you can show a supplier exactly what you bought from them last year across every site and entity, you negotiate from a different position than if you can show one site's figures. This is the most immediate return and it needs no software beyond what you own.
Consolidation
The Hackett Group's leading procurement organisations use 3.6 times fewer suppliers per billion of spend than their peers, while influencing about 20% more spend. You cannot consolidate what you cannot count.
Time back
Hackett also found analysts in leading teams spend 26% more of their time analysing data rather than collecting it. That difference is almost entirely a data quality difference.
Risk answers that hold up
RapidRatings found only 15% of enterprises fully use supplier financial health data when setting payment terms, and 30% do not use it at all — while 82% had suffered a material supplier disruption in the previous year. Knowing which suppliers matter is a data question before it is a risk question.
Six tests of whether your data is good enough
You do not need an audit to find out. Try these.
- Can someone tell you your top ten suppliers by spend for last year, with totals, within an hour?
- Does each supplier appear exactly once in your master file? Check for the same company under trading name, legal name and an abbreviation.
- Can you list every contract that expires in the next six months, with its notice period?
- Can you say what share of your spend is covered by a contract or preferred supplier?
- Can you find the same item bought by two different sites, and compare what each paid?
- Is any of this knowledge only in one person's head or on one person's laptop?
A no to questions one, two or three is common and fixable within weeks. A no to six is the one to worry about, because it is a risk that walks out of the building at five o'clock.
Why procurement data goes bad
Not through neglect exactly. Through a few structural causes that are worth naming, because each has a different fix.
- It is created at the point of least care
- Supplier records are usually set up in a hurry so an urgent invoice can be paid. Whoever types them has no stake in whether the category code is right. Fix: make onboarding a short, structured gate rather than a favour.
- Nobody owns it
- Finance thinks procurement owns supplier data, procurement thinks finance does, and IT owns the system but not the content. Fix: name one person, even part-time.
- There is no agreed taxonomy
- Without a category list everyone uses, the same spend is described five ways and cannot be added up. Fix: agree a simple taxonomy your own team recognises before buying any analytics.
- Key fields are optional
- If a system lets you save a record without a category or a payment term, some records will not have one. Fix: make the fields you actually use mandatory.
- Acquisitions and system changes
- Every merger and every migration brings another set of records with different conventions. Fix: treat data merging as part of the integration plan, not an afterthought.
The order to fix them in
Do not try to fix all four datasets at once. This order works because each step makes the next easier.
| Order | What to fix | Roughly how long | Why it comes here |
|---|---|---|---|
| 1 | Supplier master: deduplicate, standardise names, mark inactive suppliers | 2–4 weeks | Every other dataset joins to this one |
| 2 | Spend classification: categorise twelve months of transactions | 3–6 weeks | Turns transactions into something you can reason about |
| 3 | Contract register: end dates, notice periods, owners, prices | 2–4 weeks | Highest immediate return through renewals caught in time |
| 4 | Items and catalogue: descriptions and part numbers for repeat purchases | Ongoing | Slowest, and only pays back once the first three exist |
Both steps have their own detailed guides on this site: one on building a spend taxonomy a CFO will trust, and one on fixing duplicate vendor records properly.
Keeping it clean afterwards
A cleanse without a routine decays. Within a year you are back where you started, having paid twice.
A governance routine that works at mid-market scale is small:
- One named owner for supplier and spend data, even if it is part of a wider role.
- Onboarding as a gate: no new supplier record without the required fields completed.
- Thirty minutes a month checking new suppliers for duplicates against existing records.
- A one-page data standard: how names are written, which fields are mandatory, which category list is used.
- A quarterly check that the contract register still matches reality.
That is the whole programme. It fits in a few hours a month and protects everything built on top of it.
The honest part
This work has no launch event. Nobody presents a deduplicated supplier master to the board. It is the first budget line cut and the last one credited.
It is also the cheapest item on any procurement improvement list, and the one that decides whether the expensive items work. A tool bought on top of bad data produces confident, wrong answers faster. A negotiation run on partial spend data leaves money on the table you never knew existed.
If you do one thing after reading this: find out how many times your largest supplier appears in your supplier master. The answer usually starts a more useful conversation than any strategy document.
Common questions
Why is procurement data important?
Because it decides what the function can do. Four datasets — spend, supplier master, contracts, and items — sit underneath consolidation, negotiation, risk assessment, renewals and automation. When they are poor, every initiative costs more and many quietly fail. Gartner found 74% of procurement leaders say their data is not AI-ready, and Deloitte's CPOs named data quality their top internal risk in deploying generative AI at 43.97%.
What are the four key procurement datasets?
Spend (what you bought, from whom, when, in what category); the supplier master (one record per company with legal entity, bank details and status); contracts (prices, terms, end dates, notice periods, obligations); and items or catalogue data (descriptions, specifications, part numbers). Each unlocks different commercial outcomes, and they join to each other through the supplier record.
How do I know if my procurement data is good enough?
Six tests. Can someone give you your top ten suppliers by spend within an hour? Does each supplier appear exactly once? Can you list contracts expiring in six months with notice periods? Do you know what share of spend is on contract? Can you compare what two sites paid for the same item? And is any of this knowledge only in one person's head? A no to the last one is the most serious.
What should we fix first?
The supplier master, because every other dataset joins to it — typically two to four weeks to deduplicate, standardise names and mark inactive suppliers. Then spend classification, three to six weeks. Then a contract register with end dates and notice periods, two to four weeks, which usually gives the fastest financial return. Item and catalogue data last, as an ongoing effort.
Why does procurement data go bad?
Five structural causes. It is created at the point of least care, usually to get an urgent invoice paid. Nobody clearly owns it between finance, procurement and IT. There is no agreed category taxonomy, so the same spend is described several ways. Key fields are optional in the system. And acquisitions or migrations bring in records with different conventions.
How much does poor procurement data cost?
Rarely as one dramatic failure — more as a tax on everything. Analyses take longer, negotiations start from a weaker position, and risk questions get approximate answers. Hackett found analysts in leading teams spend 26% more of their time analysing data rather than collecting it, which is largely a data quality difference. Ardent puts average spend under management at 71%, so roughly three pounds in ten are spent unseen.
Do we need software to fix procurement data?
Not to start. The supplier deduplication and spend classification that deliver most of the value are data exercises, not software purchases. Buying analytics software before agreeing a taxonomy usually means paying to visualise data nobody trusts. Fix the data first, then choose tools against what you now know you need.
How do you keep procurement data clean?
A small routine beats a big project. One named owner for supplier and spend data. Onboarding as a gate, with required fields completed before a record is created. Thirty minutes a month checking new suppliers against existing records for duplicates. A one-page data standard covering naming, mandatory fields and the category list. And a quarterly check that the contract register matches reality.
How does data quality affect supplier risk management?
Directly. You cannot assess risk on a list where one company appears several times, because no single record shows your true exposure. RapidRatings found only 15% of enterprises fully use supplier financial health data in payment-terms decisions and 30% do not use it at all, while 82% had a material supplier disruption in the previous year. McKinsey found 95% can see tier-one risk but only 42% see beyond it.
What is the quickest way to show the value of better data?
Count how many times your largest supplier appears in the supplier master, then produce a single consolidated figure for what the company spent with them last year across all sites and entities. That one number usually changes how the next negotiation goes, and it makes the case for the rest of the work without a business case document.
Sources
- Gartner, 2025 Leadership Vision for Chief Procurement Officers, via Art of Procurement, 74% of procurement leaders say their data is not AI-ready.
- Deloitte, 2025 Global Chief Procurement Officer Survey, Top internal risks for GenAI (Figure 18, p.16): data quality 43.97%, governance/security/compliance 35.41%.
- 2026 Annual ProcureCon CPO Report, via Icertis, Data quality and cross-system integration cited as a barrier by 54%.
- Ardent Partners, The Metrics that Matter in 2025 (Part One) — CPO Rising, 20 October 2025. Spend under management about 71%.
- McKinsey & Company, Supply chain risk pulse 2025, 2 December 2025, 100 companies. 95% tier-one visibility versus 42% beyond tier one.
- The Hackett Group, 2025 Digital World Class Procurement research, 14 July 2025. Analysts spend 26% more time on analysis rather than manual data collection.
- The Hackett Group, What's the Digital World Class Procurement Advantage?, 24 October 2023. 3.6× fewer suppliers per $bn of spend; 20% more spend influenced.
- RapidRatings, Annual Risk Report 2026, 2 March 2026. 15% fully integrate supplier financial health into payment terms, 30% not at all; 82% material supplier disruption.
Want this run for you?
We take on the transactional half of procurement — invoices, purchase orders, supplier data and indirect spend — inside your own systems and under your approval rules. Start with a free spend audit: we measure your volumes, cycle times and exception rates, and the report is yours whether or not you go further.
Book a free spend audit


