Skip to content

Research note 1

Product matching on the hardest lines in customer requests

On our internal challenge set of the hardest lines in real customer requests, the October 2026 version of ERP Agent picks the right product for 94% of lines in 71 seconds per request. On the same base model, the March 2026 version got 68% right in 251 seconds, so the engineering since March has cut errors by 82%.

On this page

ERP Agent reads customer requests and finds the matching products in a wholesaler’s ERP catalogue. To measure progress we keep an internal challenge set built from the hardest lines in real customer requests, and we ran three versions of the agent on it, from March, June and October 2026. All three ran on the same base model, so the differences come from the engineering around the model. Because the challenge set contains only hard lines, its scores are not an estimate of average accuracy.

Scatter chart of product-matching accuracy against processing time per request. On the October 2026 base model: March 2026 68% at 251 s, June 2026 87% at 144 s, October 2026 94% at 71 s. With the model each shipped with: March 65% at 149 s, June 69% at 139 s.
Accuracy against average processing time per request on the challenge set. Orange: earlier versions on the October 2026 base model. Grey: earlier versions with the model they shipped with. Purple: the October 2026 version. Open full size

Results on the same base model

Results on the same base model
VersionRight productWrong or missingTime per request
March 202668%32%251 s
June 202687%13%144 s
October 202694%6%71 s

On the same model, the October version makes 82% fewer errors than the March version and less than half as many as the June version. It handles a request in 71 seconds, against 251 seconds for March and 144 for June.

Models and engineering

Each earlier version first ran on the model available at the time, shown in grey in the chart. Running the March code on the October base model raises it from 65% to 68%. On that same model, the October version reaches 94%, so for the March code the newer model is worth 3 points and the engineering since March is worth 26.

How we measured

Before any version ran, we fixed the correct product for every line from what sales reps actually chose, and left out lines without one clear answer. Every version received identical requests and the same organisation guide, with learning from customer data switched off, so no version could rely on answers it had seen before. We scored accuracy per product line and averaged processing time per request.

What changed

The base model is the same in all three runs, so the gains come from the engineering around it. Since March we have rebuilt how the agent retrieves candidate products from the catalogue, what context it gives the model at each step and how it reaches the final decision for every line. The October version also makes fewer model calls per request, which accounts for most of the drop in processing time.

Method notes

The figures are challenge-set results and do not describe the agent’s average accuracy across all product lines. Each result comes from one run per version, except June on its original model, which is the mean of two runs. The March version needed execution-only fixes to run in our test setup, with its matching logic and search limits unchanged.

Let’s look at your quote workflow

Test ERP Agent on your own requests

Book a call and send a few real requests, difficult ones included. You will see line by line which products the agent proposes from your catalogue.

  • Your requests and product catalogue
  • Your pricing and review process
  • A clear scope for the next step
Lauri Pelkonen

Lauri Pelkonen

Founder & CEO

Book a 30-minute call Request a demo by email

Or use the address directly