Spend Analysis

Spend Classification: Building a Taxonomy a CFO Will Trust

Classification is the part of spend analysis that decides whether anyone acts on it. Get the taxonomy wrong and you produce a defensible report that changes nothing. Here is how to design one for decisions rather than for completeness.

7 min read
Laptop displaying data analytics charts and graphs
Photo via Pexels

There is a specific way the first spend analysis review fails. The deck goes up, the categories look reasonable, and then someone senior says: that facilities number cannot be right, we spent more than that on cleaning alone. They are usually correct about their own area, and from that moment the meeting is about data quality rather than about spend.

The technical work was probably fine. What failed was that nobody could say how accurate the classification was, so one visible error was indistinguishable from systemic unreliability. That is a design problem, and it is fixable before the first slide is written.

Design for the decision, not for the standard

The first instinct is to adopt an established scheme — UNSPSC or similar — because it is comprehensive and neutral. For most mid-market organisations that is a mistake, and the reason is worth being precise about.

A standard taxonomy is designed to classify everything anyone might buy. Yours needs to classify what you actually buy, at the level at which somebody could do something about it. Those produce very different structures. A universal scheme will give you four hundred categories, of which three hundred and forty are empty, twelve hold 80% of the value, and none maps to how your organisation is actually managed.

Three levels is almost always enough. Level 1 for the board view, Level 2 for the category owner, Level 3 for the sourcing exercise. A fourth level is occasionally justified in one large category and is usually a sign that someone is classifying for its own sake.

The cleansing that has to happen first

Classification accuracy is capped by supplier data quality, and the cap binds much lower than people expect. Four problems account for most of it.

Duplicate suppliers
The same company under three records — an abbreviation, a legal name, a typo. Their spend is split three ways and each fragment lands below whatever threshold gets attention.
Parent-child relationships
Six subsidiaries of one group, each looking like a mid-sized supplier. Aggregate them and you have a top-ten relationship you did not know you had, and considerably more negotiating leverage than you thought.
Conglomerate suppliers
A supplier who sells you three unrelated things. Classifying by supplier puts all of it in one category and makes both categories wrong.
Uninformative descriptions
Line text reading "services", "miscellaneous" or a project code. Frequently 20–30% of lines, and no algorithm recovers meaning that was never recorded.

The first two are mechanical and worth doing properly, because they change conclusions rather than tidy data. The third means accepting that some suppliers must be classified at line level. The fourth is the one that sets your realistic accuracy ceiling, and it needs to be stated up front rather than discovered in the review.

Classify by value, not by line count

Automated classification handles the bulk cheaply and is confidently wrong on the residue. The mistake is to measure its success by the share of lines classified, because spend distributions are heavily skewed: a rule set can classify 85% of lines and still leave a large share of value unresolved, since the highest-value lines are frequently the most bespoke and the least well described.

A sequence that produces a defensible result:

  1. Auto-classify by supplier where the supplier sells one thing. This is fast and reliable.
  2. Auto-classify by line description where descriptions are structured — catalogue items and contracted goods usually are.
  3. Manually classify every remaining line above a value threshold. Set the threshold so that manual review covers at least 90% of the unresolved value, not a fixed number of lines.
  4. Bucket the rest as unclassified and report the number. Do not hide it in "Other".

That last point is the one that determines whether the analysis survives its first meeting. A visible, honest unclassified figure of 4% is a credibility asset. The same 4% quietly folded into "Other" is the thing that gets discovered, and when it does it takes the rest of the numbers with it.

Prove the accuracy before you present

Very few spend analyses state their own error rate, which is why they are so easily challenged. Measuring it is a day of work.

Take a random sample — a few hundred lines, weighted by value rather than uniformly, since a misclassified large line matters more. Have someone who did not build the rules classify them independently. Compare. Report the agreement rate at each taxonomy level.

Two things follow. You can now open the review with a stated accuracy rather than an implied one, which changes the meeting from an interrogation into a discussion. And the disagreements themselves are the improvement backlog: they cluster, and the clusters point at the specific rules or suppliers that need attention.

It decays, and the decay rate is predictable

A classified spend set is accurate on the day it is built and degrades from then on, because new suppliers arrive unclassified and existing ones change what they sell. A one-off exercise is therefore a snapshot with a short shelf life, and the second-year version usually costs as much as the first because it is redone from scratch.

What prevents that is small and continuous: classify new suppliers at the point of creation, refresh monthly rather than annually, and keep the rule set as a living asset rather than a project artefact. Handled that way the refresh is hours; handled as a project it is a re-run.

That is also the honest reason this work so often sits outside the internal team. Not the analysis — the analysis is the interesting part and your people should own it. The cleansing, the manual classification of the residue, the monthly refresh and the accuracy sampling are high-volume, repetitive and permanently less urgent than whatever else is happening.

Common questions

Should we use UNSPSC or build our own spend taxonomy?

For most mid-market organisations, build your own. A universal scheme classifies everything anyone might buy, which typically yields hundreds of categories where most are empty and none maps to how your organisation is managed. Design instead for the decision: a category should be something one named person could be given and asked to act on. Three levels is almost always sufficient.

What limits spend classification accuracy?

Supplier data quality, mainly. Duplicate supplier records split one company's spend across several fragments; unmapped parent-child relationships hide large group relationships; conglomerate suppliers cannot be classified at supplier level at all; and uninformative line descriptions — often 20–30% of lines — contain no meaning to recover. The last of these sets your realistic accuracy ceiling and should be stated before the analysis is presented.

How do you measure the accuracy of a spend classification?

Take a random sample of a few hundred lines weighted by value, have someone independent of the rule-building classify them, and report the agreement rate at each taxonomy level. It costs about a day. It lets you open the review with a stated accuracy rather than an implied one, and the disagreements cluster into a ready-made improvement backlog.

What should be done with lines that cannot be classified?

Report them as unclassified with the value attached. Do not fold them into an 'Other' category. A visible 4% unclassified figure is a credibility asset; the same 4% hidden is what gets discovered in the meeting, and it discredits every other number on the slide.

How often should spend classification be refreshed?

Monthly, as a small continuous task, with new suppliers classified at the point of creation. A classified dataset degrades from the day it is built as new suppliers arrive and existing ones change what they sell. Treated as an annual project it is re-run from scratch each time at close to the original cost; treated as maintenance the refresh takes hours.

Want this run for you?

We take on the transactional half of procurement — invoices, purchase orders, supplier data and indirect spend — inside your own systems and under your approval rules. Start with a free spend audit: we measure your volumes, cycle times and exception rates, and the report is yours whether or not you go further.

Book a free spend audit