Skip to content

Buyer’s guide

AI quote agent evaluation: how to pilot one on 50 of your own RFQs

Evaluate an AI quote agent on your own requests, not the vendor’s demo. Take 50 real RFQs your team has already quoted, let the agent draft them and score every line: right product, right price, edited or not, and rep time per line. Set the pass marks first. ERP Agent is made for this test: every line shows the customer’s words, the product, a 0–100 match score and the reasoning.

On this page

The checklist in brief

  • Test on 50 real requests you have already quoted, chosen by you, difficult ones included.
  • Build the answer key from what you actually sent, and keep it to yourself.
  • Score every line: product, quantity and unit, price, flagged, edited.
  • Count the lines that were wrong but looked certain.
  • Compare the rep’s review time with your manual minutes per line.
  • Re-run corrected requests to see whether the tool learns from your team.
  • Ask every vendor the same 20 questions, in writing.

Definition

AI quote agent

An AI quote agent is software that reads a customer’s request for quotation as written (an email, PDF or spreadsheet), identifies the matching products in the distributor’s own catalogue, applies that customer’s prices and prepares a draft quote, which a person reviews before it goes to the ERP, CRM or other system. See also RFQ vs RFP vs quotation.

Why isn’t a vendor demo enough to evaluate an AI quote agent?

In a demo, the vendor chooses the request, and it is usually clean. Reading text off a PDF is not where quoting time goes. Your reps’ time goes on what comes next: which of your products “SS flat 40x5” or a customer’s own article code means, what this customer pays for it, and whether a line is doubtful enough to check.

That work is measured per line. According to our case study, a sales rep spends on average 4–7 minutes per line when creating quotes and orders by hand. A tool that only reads the document leaves that work with the rep. ERP Agent is built for the part that takes the time: it identifies the right product in your own catalogue from the customer’s own words, brings in that customer’s pricing and shows on every line how sure it is. A pilot on your own requests puts that in your numbers. See the cost of manual quoting per line.

RFQ automation pilot checklist: seven steps

Run the same steps with every vendor on the shortlist, so the results are comparable.

  1. Decide what you are testing

    Quotes, sales orders or both; which branches and customers; which reps review the drafts. Write the pass marks down before anyone sees a draft.

  2. Pick 50 requests

    Recent ones, exactly as they arrived: original emails with attachments, not retyped lists. Weight the mix towards what fills your reps’ days, and keep a few they dread.

  3. Build the answer key

    Take each quote you sent from the ERP: product, quantity, unit and price per line. Correct known mistakes. Where two products were both acceptable, record both.

  4. Give the vendor production inputs

    The catalogue export, the price data or ERP connection you would use live, and the raw requests, but not the answer key. Keep the 50 test requests out of any quote history the tool learns from, so the result shows how it handles requests it has never seen.

  5. Let your reps review the drafts

    Reps correct each draft in the vendor’s review screen as they would before sending, and note the time. Fill in the scoring sheet yourselves rather than relying on the vendor’s report.

  6. Re-run corrected requests

    Once the reps have fixed the wrong lines, send 10 of the corrected requests again, or new ones that ask for the same products in the same words. A tool that learns from your team now picks the corrected product; one that does not repeats the mistake every time.

  7. Decide against the pass marks

    Calculate the numbers below for each vendor, then read the failed lines by type. Some failures are fixed in setup; others show how the tool works.

Which requests should go in the 50-request sample?

An example split for a wholesaler that receives most enquiries by email. Adjust it to your own inbox.

Which requests should go in the 50-request sample?
Request typeExample countWhy it is in the test
Short everyday emails in free text15Most of the volume. The tool must beat a rep here, or nobody will use it.
Long lists in Excel or CSV10Length, repeated lines, units, and how a big request is handled.
PDFs, including a poor scan8Tests reading as well as identifying the products.
The customer’s own codes or another brand’s part numbers7Whether the tool can use your customer part numbers and cross-references.
Vague lines: no grade, no size, “same as last time”5Whether the tool flags doubt or guesses.
Another language, a forwarded chain, a photo5Requests that arrive anyway. A tool that reads only clean documents fails here.

Example counts, not a rule. Around 50 covers each type several times and can still be scored by hand.

The scoring sheet: one row per request line

Copy these columns into a spreadsheet, one row for every line of every request.

The scoring sheet: one row per request line
ColumnWhat to record
Request and lineThe line exactly as the customer wrote it.
Expected productFrom the answer key, with any acceptable alternative.
Proposed productWhat the tool put on the draft, or “none”.
Right productRight, acceptable alternative, wrong, or no proposal.
Right quantity and unitYes or no. Watch metres against lengths and packs against pieces.
Right priceRight, wrong or not priced, against what this customer pays under your terms.
FlaggedDid the tool mark the line as uncertain?
EditedWhat the rep changed: product, quantity, price, or a line added or removed.
Rep timeMinutes from opening the draft to approving it, per request.
Failure typeProduct data, customer code, description, price source, or file not read.

Seven numbers that decide the pilot

Seven numbers that decide the pilot
NumberHow to calculate itWhat it tells you
Right productRight or acceptable lines ÷ all linesHow much of the product work the tool takes off your reps.
Right priceCorrectly priced lines ÷ lines with the right productWhether it applies your customer pricing or leaves it to the rep.
Lines needing editsLines the rep changed ÷ all linesThe work left for your team.
Confident and wrongWrong lines not flagged ÷ all wrong lines (a wrong line has the wrong product, quantity, unit or price, or no price)How often a rep must catch an error with no warning.
Rep time per lineTotal rep minutes ÷ linesThe saving against your manual baseline.
Failed requestsRequests the tool could not read or draftFormats you would still handle by hand.
Fixes keptRe-run lines now right ÷ re-run lines your reps had correctedWhether your team’s knowledge stays in the tool.

Report each number per request type as well as overall. An average can hide a tool that handles short emails well and long lists badly.

Example: five scored lines

A worked example: one request from a building services contractor, scored line by line.

Example: five scored lines
Line as writtenProposed productProductQty and unitPriceFlagged
10 x ball valve DN25 PN40 full boreBall valve DN25 PN40, full boreRightRightRightNo
60 m copper pipe 22mmCopper tube 22 × 0.9 mm, 3 m lengthRightWrong: 60 lengths, not 20RightNo
20 x our ref 44-118Pipe clip 54 mm, rubber linedAcceptable alternativeRightRightYes
12 x spiral duct Ø250Spiral duct Ø200Wrong: Ø250 askedRightRightNo
6 x radiator valve angle 1/2 inchRadiator valve, angle pattern, DN15RightRightNot pricedYes

The sheet reads as 4 of 5 right products, 3 of those 4 correctly priced and 3 of 5 lines edited. Three lines were wrong (the tube quantity, the duct size and the missing valve price), and only the unpriced valve was flagged: 2 of 3 wrong lines were confident and wrong. The 180 m of tube, against 60 m asked, is the error the customer would see.

Watch the lines that were wrong but looked certain

Important

A flagged wrong line costs the rep a look. An unflagged one goes to the customer, and after a few of those, reps check every line again and the saving is gone. Ask each vendor how its tool shows doubt, and count how often it stayed quiet.

Where ERP Agent catches these errors

The errors in the example are the ones ERP Agent is built to stop before a customer sees them:

  • Doubt you can see. Every line carries a match score from 0 to 100, and lines under 80 are marked for review. Open a line to read the agent’s reasoning or take an alternative product.
  • Sizes and ratings checked. ERP Agent checks the size, voltage or rating class a line states and pushes down products that contradict it, such as a Ø200 duct when Ø250 was asked.
  • Units that add up. For steel and metal products, it converts between lengths, metres, kilos and pieces and recognises standard lengths.
  • This customer’s prices. Customer pricing comes from the data and rules configured for your company; once your ERP is connected, from its own price logic, customer discounts included. The rep can edit any price before sending.
  • No guesses. A line it cannot match goes on your placeholder code, marked for the rep, instead of on a near match.

How do you measure rep time per line fairly?

Measure the manual baseline first, on similar requests: from opening the email to the quote saved in the ERP, including look-ups, price checks and keying. Time an experienced rep and a newer one.

In the pilot, time the rep from opening the draft to approving it. Count the tool’s processing time separately: it runs in the background while the rep works on something else. ERP Agent drafts a typical request in 3–8 minutes and a request of several hundred lines in about half an hour; fast mode handles requests of up to about 50 lines in 1–3 minutes. At the case study’s 4–7 minutes per line, several hundred lines take a rep well over a working day by hand.

Score each rep’s first few drafts separately while they learn the screen. The monthly saving is quote lines per month × (manual minutes per line − pilot minutes per line) ÷ 60 × loaded hourly cost. Enter the difference as minutes per line in the quote cost calculator to see the hours and cost.

What pass marks should you set before the pilot?

Write them down and share them with every vendor before the first draft:

  • A minimum share of right products on everyday requests, and a separate one for long lists and customer codes.
  • A ceiling on confident and wrong lines, by product group.
  • A maximum rep time per line, below your manual baseline.
  • Must-pass requests: your largest customers’ formats.
  • A minimum share of fixes kept when corrected requests are re-run.
  • What must work on day one and what follows once the ERP connection is live, such as write-back.

Benchmark the tool against your own reps

Give 10 of the 50 requests to two experienced reps separately. Where they chose different products, both answers go in the answer key: no tool should be held to a stricter standard than your team. The comparison also shows where your product data is ambiguous.

How to evaluate quote automation software: 20 questions for every vendor

Send the same list to each vendor before the pilot and ask for written answers. The right-hand column describes a strong answer.

How to evaluate quote automation software: 20 questions for every vendor
#TopicQuestionA strong answer
1ERP connectionDoes it work with our ERP, CRM or other system?Yes, whatever you run: the vendor sets up the connection for your company during onboarding, with a named plan.
2ERP connectionWhat exactly is written back, and as what?A quote or a sales order created in your ERP, with the fields listed in writing: customer, delivery address, lines, units, prices and references.
3ERP connectionWhat happens if a product code or customer on the draft is not in the ERP?The write stops, the rep is told why, and nothing is saved halfway.
4ERP connectionHow soon can we start, and do we need a new system?The same week, on top of the ERP you already run: a catalogue file is enough to start, and the connection follows. No replacement project.
5RequestsWhich formats does it read as they arrive?Email bodies and attachments (PDF including scans, Excel, CSV, Word, images, forwarded emails), with no template per customer format.
6ProductsHow does it decide which product a line means from a description, the customer’s code or another brand’s part number?The sources it uses (your catalogue, past quotes, customer codes, cross-references), shown on your own requests, with the customer’s words kept next to each match.
7ProductsDoes it check the sizes and ratings a line states?Products that contradict a stated size, voltage or rating class are pushed down, and the reasoning says why.
8ProductsCan it read a request in one language against a catalogue in another?Yes, with the original line kept next to the proposed product.
9ProductsWhat happens to a line it cannot match?It goes on your placeholder code and is marked, never on a near match without warning.
10PricingWhere do the prices on the draft come from?One named source for every price: your ERP’s own price logic or the price rules agreed for your company.
11PricingHow are customer-specific prices applied, and what if the customer is not identified?The customer is found from the request, with a visible warning when none matched.
12PricingCan the rep change every price before anything is sent?Yes, on the draft before sending, not only in the ERP afterwards.
13ReviewCan anything reach the ERP or a customer without a person approving it?No. A person reviews by default before anything is written to the ERP, and nothing goes to a customer without approval.
14ReviewWhat does the rep see on each line?The request text beside the proposed product, a score, the reasoning and alternatives to pick.
15ReviewCan the rep change the draft without retyping it?Rows added or removed and quantities or prices changed on request, each change approved line by line.
16ReviewHow does it learn from the lines our reps fix?From the lines your team sends: the next time a customer writes the same thing, it picks the same product.
17Data and priceIs the service GDPR-compliant, and what do we agree before sharing data?GDPR-compliant, with the data the tool needs and the systems it can access agreed before setup.
18Data and priceDo the AI models retain our data or train on it?A clear written answer, such as zero data retention.
19Data and priceHow is our data kept apart from other customers’ data?Access keys scoped to your organisation only, so no other tenant’s data can be reached.
20Data and priceHow is it priced after the pilot?The pricing model in writing and an itemised quote before you start.

ERP Agent’s answers to the 20 questions

Score our drafts with the same sheet. Here is what you will find.

ERP Agent’s answers to the 20 questions
TopicERP Agent
ERP connectionERP Agent integrates with Business Central, SAP, NetSuite, Visma, IFS, Sage and any other ERP or CRM. We set up the connection for your company during onboarding, and you can start with a catalogue file the same week. It works on top of your ERP: no replacement project.
Write-backCreates the quote or sales order in your ERP, straight from the reviewed draft. Product codes are checked against the ERP first; if one does not exist there, nothing is saved or sent. Starting from a catalogue file, the reviewed draft is delivered as a file.
RequestsReads email bodies and attachments: PDF (including scanned PDFs), Excel, CSV, Word, images and forwarded emails. Each customer company gets its own agent mailbox; reps forward requests there or paste them into the web app. Your own systems and AI agents can send requests through the REST API and the MCP server.
ProductsIdentifies the right products in your own catalogue from abbreviations, dimensions, materials, technical descriptions, part numbers and customer codes, across languages. Checks the size, voltage or rating class a line states. An unmatched line goes on your placeholder code, not on a guess.
PricingBrings in customer pricing from the data and rules configured for your company; once your ERP is connected, from its own price logic. Warns when the customer could not be identified. Reps can edit any price before sending.
ReviewA 0–100 match score on every line, lines under 80 marked, the agent’s reasoning and alternative products. Quote chat adds or removes rows and changes quantities or prices, with every edit approved line by line. A person reviews by default before anything is written to your ERP, CRM or other system.
LearningLearns from the lines your team sends: the next time a customer writes the same thing, it picks the same product.
Data and priceERP Agent is GDPR-compliant, and models run with zero data retention. Before setup, we agree which data the agent needs and which systems it can access. API keys are scoped to one organisation. Pricing is a setup fee and a monthly subscription, with an itemised quote first. See security and pricing.

At LVI-WaBeK, ERP Agent automates quote lines with over 95% accuracy, and our customers average 95% automation. Read the LVI-WaBeK story.

What are the red flags in a quote automation pilot?

  • The vendor wants to choose the test requests, or have them retyped into a template.
  • Each customer’s document format has to be mapped before the tool can read it.
  • The tool reads the document but leaves choosing the product to the rep.
  • The vendor asks for the answer key before the test.
  • Pilot results come as one headline figure, with no line-level sheet to check.
  • Every line shows the same confidence, right or wrong.
  • Drafts can only be corrected in the ERP afterwards, so the tool never learns.
  • Getting started means replacing or re-implementing your ERP.

Frequently asked questions

How many RFQs do you need to evaluate an AI quote agent?

Around 50 real requests: enough to include each request type several times, and few enough to score line by line by hand. A much smaller sample says little about the awkward formats, where tools differ most.

What accuracy should an AI quote agent reach?

Set the pass mark against your own reps on your own requests, with one mark for right products and another for lines that were wrong without a warning. For reference, our customers average 95% automation, and at LVI-WaBeK, ERP Agent automates quote lines with over 95% accuracy.

What is the difference between automation rate and accuracy?

Accuracy is the share of lines the tool got right. Automation is the share of the quoting work the tool takes off your reps. The scoring sheet above measures both on your own requests: right product and right price for accuracy, lines needing edits and rep time per line for automation.

Do we need an ERP connection to run an RFQ automation pilot?

No. A catalogue export is enough to test how well the tool identifies products, and with ERP Agent you can start that way the same week. Customer pricing and write-back follow once the connection is live: ERP Agent integrates with Business Central, SAP, NetSuite, Visma, IFS, Sage and any other ERP or CRM, and we set up the connection for your company during onboarding. See how quotes reach your ERP, CRM or other system.

How long should a quote automation pilot take?

With ERP Agent you can start the same week: send a catalogue file and your 50 requests, and the drafts arrive in the background, a typical request in 3–8 minutes. Most of the pilot time is your own team scoring the drafts: 50 requests of 10 lines each means 500 lines, so book your reps’ time for it before you start.

What data do we share with a vendor for a pilot?

The catalogue, the 50 requests and price data or an ERP connection. Requests contain customer names and contact details, so agree in writing how they are processed under the GDPR. With ERP Agent this is settled before setup: ERP Agent is GDPR-compliant, models run with zero data retention, and we agree which data the agent needs and which systems it can access.

Can we run this pilot with ERP Agent?

Yes. Bring real requests from your inbox, including the difficult ones, start from a catalogue file and score our draft lines with this sheet, line by line. Book a call to agree the scope.

Let’s look at your quote workflow

Score our drafts with this sheet

Bring real requests, including the difficult ones, and score ERP Agent’s draft lines with this sheet. You can start from a catalogue file the same week.

  • Your requests and product catalogue
  • Your pricing and review process
  • A clear scope for the next step
Lauri Pelkonen

Lauri Pelkonen

Founder & CEO

Book a 30-minute call Request a demo by email

Or use the address directly