On this page
ERP Agent reads customer requests and finds the matching products in a wholesaler’s ERP catalogue. To measure progress we keep an internal challenge set built from the hardest lines in real customer requests, and we ran three versions of the agent on it, from March, June and October 2026. All three ran on the same base model, so the differences come from the engineering around the model. Because the challenge set contains only hard lines, its scores are not an estimate of average accuracy.

Results on the same base model
| Version | Right product | Wrong or missing | Time per request |
|---|---|---|---|
| March 2026 | 68% | 32% | 251 s |
| June 2026 | 87% | 13% | 144 s |
| October 2026 | 94% | 6% | 71 s |
On the same model, the October version makes 82% fewer errors than the March version and less than half as many as the June version. It handles a request in 71 seconds, against 251 seconds for March and 144 for June.
Models and engineering
Each earlier version first ran on the model available at the time, shown in grey in the chart. Running the March code on the October base model raises it from 65% to 68%. On that same model, the October version reaches 94%, so for the March code the newer model is worth 3 points and the engineering since March is worth 26.
How we measured
Before any version ran, we fixed the correct product for every line from what sales reps actually chose, and left out lines without one clear answer. Every version received identical requests and the same organisation guide, with learning from customer data switched off, so no version could rely on answers it had seen before. We scored accuracy per product line and averaged processing time per request.
What changed
The base model is the same in all three runs, so the gains come from the engineering around it. Since March we have rebuilt how the agent retrieves candidate products from the catalogue, what context it gives the model at each step and how it reaches the final decision for every line. The October version also makes fewer model calls per request, which accounts for most of the drop in processing time.
Method notes
The figures are challenge-set results and do not describe the agent’s average accuracy across all product lines. Each result comes from one run per version, except June on its original model, which is the mean of two runs. The March version needed execution-only fixes to run in our test setup, with its matching logic and search limits unchanged.
