Your Spend Data Is Lying to You — Here's the Fix
AI spend classification builds a clean, consistent taxonomy across millions of transactions, exposing maverick spend and the savings actually within reach.
Every procurement leader I've met believes their spend cube is roughly right. It almost never is. Twenty percent of it lands in "miscellaneous," the same supplier shows up under six spellings, and half the "IT" spend is actually facilities. You can't manage what you've mis-tagged.
Classification has always been the grunt work of spend analysis. Someone maps transactions to categories by hand, it takes weeks, it's stale the day it's finished, and the moment a new vendor appears the whole thing drifts again.
Where AI changes the economics
Machine-learning classifiers can categorize transactions at scale by reading the supplier name, the line description, the GL account, and the context around each purchase. GenAI helps with the ambiguous long tail — the free-text descriptions that a rules engine can't parse ("misc. parts for line 3"). Together they'll classify millions of transactions to a UNSPSC-style taxonomy in hours, with consistency a manual team can't match across that volume.
The payoff isn't a prettier chart. It's that you can finally see addressable spend. Once categories are clean, maverick spend — purchases made off-contract or outside preferred suppliers — stops hiding. We routinely find 15–30% of a category running maverick that leadership swore was under control. That's the exact spend where negotiated rates get left on the table.
From there the savings math gets real. Consolidate fragmented spend under fewer suppliers, move off-contract volume onto negotiated agreements, and spot tail suppliers you're paying premium rates to for no reason. Clean classification is what turns "we think there's savings" into "here's the addressable number and the three categories to go after first."
The parts people underestimate
A taxonomy is a decision, not a default. If you don't align categories to how your organization actually buys and how sourcing is structured, you'll get technically-correct classifications that nobody can act on. Spend the time here — it's the difference between an interesting dataset and a usable one.
Classification confidence varies wildly by category. Direct materials with clean part numbers classify beautifully; professional services and one-off free-text spend are messier. So we keep a human review loop on low-confidence and high-value transactions rather than trusting the model blindly on a $2M consulting invoice.
And this isn't set-and-forget. New suppliers and new spend patterns appear constantly, so the model needs feedback — a mechanism where a category manager corrects a misclassification and the system learns from it. Skip that and accuracy decays quietly over a couple of quarters.
One more thing worth saying plainly: the classification is only as good as the source data. Missing descriptions, blank GL codes, and duplicate vendor records all cap what any model can do. Often the first win isn't the AI at all — it's the vendor master cleanup the AI made you finally confront.
Get it right, though, and every sourcing conversation afterward starts from facts instead of hunches.
