Featured answer: AI can reliably automate the parts of sourcing that are rules-based and already documented: matching a quote against a spec, normalizing supplier data, flagging a missing certificate. It cannot automate the two questions that decide whether an order is worth placing at all: whether the supplier is real, and whether the price is honest. Those require someone on the ground. Judge any sourcing AI tool by which of those two categories it touches first.
Every few months a new procurement tool claims it will run your sourcing. The pitch is always the same: your team spends half its time on manual work, an agent can do that work, therefore you get the savings. At a 2026 supply-chain conference, Tim Spencer, co-founder of the AI procurement company Didero, put a number on the upside and then explained why it is harder than it sounds. Roughly half of a supply chain team's time goes to work that could, in principle, be automated. And yet most people are aware of agents without having used one, because physical supply chains move and change in ways that break the assumptions agents are built on [1].
That gap between the demo and the dock is the most useful thing a small buyer can think about, because for a company without an IT department the question is not whether AI is powerful. It is which specific part of your sourcing it can be trusted with, and which part it will quietly get wrong. There is a clean dividing line, and once you can see it you can evaluate any tool in about ten minutes.
The dividing line: rules versus judgment

The dividing line is not model quality. It is whether the task has a written rule that anyone could apply identically twice.
On the automatable side, the work is genuinely mechanical. A quote arrives in three formats from three suppliers and needs to become one comparable table; that is a parsing problem with a known answer. A product specification has a required attribute list, and a supplier's response either covers each attribute or does not; that is a checklist. A shipment is missing a certificate that the same product always required last quarter; that is a comparison against a known requirement. A purchase order, an invoice, and a packing list disagree on a quantity; that is arithmetic. All of these have something in common: the correct answer is written down somewhere, and a machine that can read both sides will reach the same answer a careful analyst would.
On the other side, the work that determines whether an order is worth placing has no rule to apply. Is this factory a factory or a trading company with a workshop out back? Does this quote contain the cost the supplier will absorb silently and then recover on the next order? Will this factory actually run your tooling on the date it promised? Is the person on the other end of WeChat answering for the entity named on the invoice? None of these have a written rule. There is no spec to check them against, which is exactly why they keep being done by people who visit.
| Sourcing task | Automatable | Why |
|---|---|---|
| Normalizing quotes from different formats into one comparison | Yes | The rule exists: same units, same currency, same line items |
| Checking a quote against a written specification | Yes | The spec is the rule, and both sides are documents |
| Flagging a missing certificate the product always required | Yes | The requirement is historical and documented |
| Matching PO, invoice, and packing list quantities | Yes | Arithmetic on three known values |
| Building a supplier database from filings and public records | Yes | The sources are structured and public |
| Telling a real factory from a trading company | No | The difference is what the workshop looks like when nobody is presenting |
| Knowing whether a quoted price hides a cost | No | The hidden cost is precisely the part that is not quoted |
| Judging whether a delivery promise is real | No | It depends on that factory's current order book, which is invisible to you |
| Deciding whether a sample result represents mass production | No | It requires knowing what changed on the line in between |
Read that table bottom-up, because the answer is not evenly distributed. Everything above the line was already expensive and slow because it was manual. Everything below it was never going to be captured, no matter how good the model gets, because the evidence is physical and the verifier has to be there.
Why the bottleneck sits in the data, not the intelligence
If you want to understand why agents underdeliver on physical supply chains, look at the paperwork the government itself deals with. In its own documentation for the single-window modernization program, US Customs and Border Protection states that 47 partner government agencies are involved in the trade process, that importers and exporters must complete nearly 200 forms, and that the existing processes remain largely paper-based and require manual entry into multiple government-owned electronic systems, so the same data is often submitted repeatedly to multiple agencies at multiple times [2].
That is the honest headline about supply chain data. It is not that a company is disorganized. It is that a cross-border transaction is genuinely spread across dozens of record systems, most of them owned by someone else, and a meaningful share of it still arrives as documents a person has to read. An agent that automates a process still needs that process to be written down somewhere. Where it is not, the agent is guessing, and a fast guess is more dangerous than a slow one.
CBP has spent years proving the other half of the argument. Its Automated Commercial Environment is the single window through which importers file manifests, entries, and entry summaries, and its own materials describe filing an entry, filing an entry summary, running reports for in-house audit, and responding to a request for documents as standard capabilities [3]. Where the rule is clear and the data is structured, this is what working automation looks like: not clever, just a rule that everyone follows. The lesson transfers directly to a buyer's own sourcing stack.
The exception is where the cost is
There is a design idea from the same conference worth taking seriously, because it is the opposite of automation and that is exactly why it works. Spencer's suggestion was an agentic execution layer that sits above the enterprise system, the warehouse system, and the transportation system, so that most transactions run without a person and humans are involved only for exceptions [1].
That is the correct shape for a physical supply chain. Most of a sourcing process really is routine once you have run it many times, and only a small part genuinely needs a person. The mistake is treating those two parts as one workstream, or assuming the routine part is the one worth automating first because it is easier to measure. In practice the routine part is where your time savings come from, and the exception part is where your losses come from. Both are worth doing, but they are not the same project.
The trap in the exception path is that an exception is by definition something your rules did not anticipate, which means the model is reasoning about something it has no precedent for. This is where trusting the output without checking it becomes expensive. A flagged discrepancy is fine to auto-resolve if the resolution is conservative. A guessed answer about a factory, an origin, or a compliance status is not, because there is no conservative fallback.
| Sourcing task | Who should own it | Failure mode if automated wrongly |
|---|---|---|
| Quote normalization and spec checking | Software, unattended | Low. Worst case is a table you have to re-read |
| Certificate tracking and document chase | Software, unattended | Low. A missing item resurfaces at inspection |
| Cost build-up reconstruction | Software, then human review | Medium. Plausible numbers that hide a real cost |
| Supplier identity verification | Human on the ground | High. You supply a counterparty you cannot describe |
| Certification and compliance verification | Human, against documents and evidence | High. Penalties attach to the importer of record |
| Negotiation and dispute handling | Human | High. It sets the terms of everything downstream |
One more column matters when you are sequencing the work, which is when the discrepancy has to be caught. The right measure of an exception process is not how many exceptions it closes, but how early.
| Discrepancy type | Caught before the PO closes | Caught after the vessel sails |
|---|---|---|
| Specification attribute mismatch | Rework the order, no cost | Inspection finding, rework freight, possible penalty exposure |
| Missing or wrong-revision certificate | Chase the supplier, ship on time | Cargo held or excluded at the border |
| Quantity mismatch between documents | Correct the paperwork, no cost | Demurrage, storage, or a claim against the wrong party |
| Supplier identity not verified in advance | Skip the supplier, no cost | Counterparty you cannot describe, unpaid or undeliverable |
| Cost structure never questioned | Accept the quote | Discovered in a repeat order, priced into your next run |
The pattern is not subtle. Almost every discrepancy that gets expensive is one that was available to catch while the purchase order was still open. That is the whole argument for putting the exception workflow in front of the automation rather than behind it, and it is also the asymmetry that runs through both tables: the tasks at the top are worth automating precisely because being wrong is cheap, while the tasks at the bottom are worth automating only when being wrong is survivable. For a sourcing operation it usually is not.
The three questions a buyer can actually verify
Claims about automation are cheap and everyone makes them. The following questions are cheap to ask and difficult to answer well, which makes them the fastest way to find out whether a tool will survive contact with a sourcing operation. None of them require technical knowledge to evaluate. They require the vendor to be specific.
1. Which of my tasks does this system refuse to do?
A good answer names specific exclusions and explains why. A tool that will not classify products under the Harmonized Tariff Schedule [5] should say so plainly, because classification is a rules problem with published rules and a human signed up for the liability. A tool that will not verify a supplier's business license should say that too, and if you want a reference point for what that verification actually involves, our guide to verifying a Chinese supplier is a real factory, not a middleman sets out what the records do and do not tell you.
The shape of the answer matters more than its content. "Our AI handles most scenarios" is not an exclusion. "We do not make origin determinations" is an exclusion, and it is the kind that tells you the vendor understands the difference between reading a document and making a declaration. When you have three tools in a row answering the first way, the pattern is the market's, not your negotiating position.
2. What happens when it does not know?
Every tool claims some confidence threshold. Almost none will tell you what is below it.
This is the question that separates a product from a demo, and it deserves a straight answer rather than a reassuring one. Ask what the system returns for an input it has no precedent for, which in sourcing is a weekly event rather than an edge case. A tool that returns a plausible answer is worse than no tool in this specific domain, because the failure is silent and you will find out at the worst possible moment. A tool that says it is out of scope and hands the task back to a person is doing the defensible thing, and it should be cheaper than the first.
There is a concrete way to test this without a technical team. Take a real quote and a real specification from your last order, and deliberately give the tool an input with something missing, something contradictory, and one attribute it has never seen. If all three come back with equal confidence, you have learned what you need to know about it.
3. Who owns a discrepancy it cannot resolve?
Discrepancies are the normal case in sourcing. Quantities that do not match, certificates that cover the wrong revision, a port of exit that changed mid-voyage. A tool designed for a world where clean inputs dominate will not have a good answer here, because this is where its assumptions were never tested.
The answer you want is a workflow: the system packages the relevant documents, routes the case to a named person, and records what was decided. That is what an execution layer looks like when it is built for this domain. If the answer is that the system improves over time, the failure mode is not covered, and no amount of pilot data changes that.
The cost of automating the wrong half
The most expensive outcome is not a tool that fails. It is a tool that works on the wrong work.
A sourcing operation that automated only its document layer would get a real reduction in hours spent on quote formatting and certificate chasing. That reduction would show up in a time study, and the time study would look convincing, because nobody measured the errors that were never caught. An order goes out with a misstated material specification because the checklist had the wrong attribute in it. A supplier certificate was chased and filed on time, for the wrong product revision. The exception workflow routes a discrepancy to a person three days after the vessel sails instead of three days before the purchase order closes.
None of those are AI failures in the sense the vendor would recognize. The system did exactly what it was configured to do. The failure is upstream: someone decided which half of the work was the problem, and that person was usually looking at the half that is easiest to count.
There is also a subtler version. Once the routine half is automated, the exception half becomes the entire job, but it arrives without the structure it used to have. When a person handled the exceptions, they accumulated judgment invisibly: they learned that this supplier always misses the coating date, that this factory's tooling record is fiction, that this broker's invoices never match the entry summary. None of that lived in a system, so none of it can be migrated into one. What an automation project actually produces in that scenario is a smaller team holding the same exceptions with less informal knowledge behind them.
The specific trap in automated extraction is provenance. A number that came out of a system is not more reliable than a number a person typed; it is a number whose origin nobody remembers. Once supplier data has flowed through extraction, it becomes very hard to tell which fields were read from a document, which fields were inferred, and which fields came from a supplier's own marketing material. When a compliance question arrives two years later, that distinction is the whole answer, and the tool that produced the record may not be able to supply it.
The honest way to evaluate any of these tools is therefore not to ask what it can do. Ask which of your two halves it touches, and whether anyone has named the exclusions. A vendor who cannot name exclusions is selling the first half as if it were the whole, which is the most common framing in this market and the least useful one.
What to automate first, and in what order

If you take nothing else from the automation pitch, take the ordering. Automate the document work first, because the rule exists and the failure is cheap. Normalize every quote before you compare any of them, because a comparison across mismatched units and currencies is not a price comparison. Build the specification into a checklist so supplier responses become completable rather than readable. Track certificates against the history of what that product has required before. These are unglamorous, they are where your own time actually goes, and they do not require you to trust anyone's judgment.
Then stop and look at what is left. What remains is the supply chain knowledge your business actually runs on, and the reason buyers pay outside for it: whether the factory is real, whether the price is honest, whether the documentation survives a customs audit. Keep a human on that. The value of your operation is that someone can go and look.
There is a practical test for whether you have reached that line. Take the last three sourcing problems you solved and ask, for each one, whether the answer came from a document or from being somewhere. If most of them came from a document, you have a documentation problem and automation will genuinely help. If most came from being somewhere, you have a field knowledge problem, and the honest conclusion is that your budget belongs in the field and your time belongs in the document layer.
Whichever half you are in, the arithmetic underneath it is the same one we walk through in our total landed cost walkthrough, and any tool you evaluate should be able to produce a number you can check against that method rather than a dashboard you cannot.
That test also tells you which vendor claims to believe. A tool that automates document work makes a modest, checkable promise. A tool that promises to replace supplier judgment is asking you to believe that the value you pay for can be reproduced from documents, and the customs record contains at least one recent, very expensive example of what happens when that assumption drives a declaration.
Then measure the right thing. Not the hours saved, which is easy and misleading, but the proportion of exceptions caught before the purchase order closes rather than after the vessel sails. That metric moves for the first time when the exception workflow exists, which is the clearest signal that a project has reached the part of the business that matters.
The question to ask the vendor, and the one to ask yourself
Ask the vendor: which of my sourcing tasks does this system refuse to do? If the answer names specific exclusions, you are talking to someone who has thought about your operation. If the answer is a list of capabilities, you are talking to a product demo. The same vendor should be equally willing to name the tasks it will never touch, because a tool that claims to cover supplier identity and origin determination is describing something that does not exist.
Then ask yourself the harder question: if this system worked perfectly, what would still be left that requires my judgment? If the honest answer is not much, your operation may not have the substance you assume it has. If the answer is a long list, that is not a reason to avoid automation. It is a reason to automate the other half with real enthusiasm and keep your budget and your attention on the part that remains.
That is the version of this conversation worth having. The half of your team's time that could be automated is real, and reclaiming it is worth money. It is also the half that was never where the money was made.
Footnotes
David Maloney, "The challenges of adopting AI in physical supply chains," The Supply Chain Xchange, October 7, 2026. https://www.thescxchange.com/tech-infrastructure/technology/the-challenges-of-adopting-ai-in-physical-supply-chains
U.S. Customs and Border Protection, "Appendix A: Automated Commercial Environment (ACE)," U.S. Department of Homeland Security, March 2020. https://www.dhs.gov/sites/default/files/publications/privacy-pia-cbp-ace003b-appendixa-march2020.pdf
U.S. Customs and Border Protection, "How to Use the Automated Commercial Environment (ACE)," accessed October 8, 2026. https://www.cbp.gov/trade/automated/how-to-use-ace
U.S. Customs and Border Protection, "ACE Transaction Details," accessed October 8, 2026. https://cbp.gov/trade/automated/ace-mandatory-use-dates
United States International Trade Commission, "Harmonized Tariff Schedule of the United States," accessed October 8, 2026. https://hts.usitc.gov/

