Every AI shopping assistant vendor has impressive numbers on their website.
3x conversion. 4x revenue per visit. 11% of total site revenue attributed to AI. Fifty percent fewer returns.
Some of those numbers are real. Many of them are constructed from metrics that sound rigorous and aren't. The difference between a vendor whose tool is genuinely driving revenue and a vendor who has configured their analytics to produce a compelling number is not visible from the outside.
This guide is for fashion brands who want to know the difference, and who want to measure their own deployment honestly, regardless of what their vendor's dashboard says.
Why most AI shopping assistant measurement is wrong
The standard measurement approach looks like this: a merchant adds an AI shopping assistant, tracks the revenue from sessions that included an AI interaction, and attributes that revenue to the tool.
The problem is that shoppers who choose to use an AI shopping assistant are not a random sample of your visitors. They're more engaged, more purchase-intent, and more likely to convert regardless of whether the AI helped them. A shopper who opens a chat and asks "help me build an outfit for my sister's wedding" was already meaningfully further along in the purchase funnel than a shopper who browsed for 45 seconds and left.
When you attribute all the revenue from AI-engaged sessions to the AI tool, you're crediting the system for sales that would have happened anyway. The actual incremental revenue, the sales that happened because the AI existed, that wouldn't have happened without it, is a subset of the attributed total.
How large a subset? It varies. Some vendors, when asked to run a proper controlled test, find that their "influenced revenue" number overstates true incremental impact by 2 to 3x. Others hold up fine. You won't know which category yours is in until you measure it correctly.
The only measurement that actually answers the question
The honest measurement of any AI shopping assistant is incremental lift over a holdout.
A holdout is a control group, a percentage of your shoppers who browse your store normally, without seeing the AI tool, during the same period as the shoppers who have access to it. After a defined period (14 to 30 days is typical), you compare outcomes between the two groups.
The difference in conversion rate, AOV, and revenue between the exposed group and the holdout is the actual increment the tool drove. Not sessions that included a chat. Not revenue that occurred in the same visit as an AI interaction. Revenue that would not have existed without the tool.
This is what statisticians call incremental lift measurement. It's the same methodology used to evaluate the true impact of any marketing or product intervention, from pricing tests to email campaigns. It's the standard for rigorous performance measurement and it's rarely used for AI tools, largely because most vendors don't offer it, and some know their numbers wouldn't survive it.
How to run a proper holdout test
You don't need a vendor to do this for you. You can run a holdout test on any AI shopping assistant deployment with the right configuration.
Step 1: Define the holdout percentage. Five to ten percent of your traffic is sufficient for statistical significance at most Shopify store traffic levels. The holdout group should be randomly assigned, not segmented by geography, device, or any other variable that would make the groups non-comparable.
Step 2: Implement the holdout cleanly. The holdout group must not see the AI tool at all, not just not use it. If the widget is visible but they don't open it, you can't meaningfully compare them to shoppers who didn't see it, because the mere presence of the widget may affect behavior.
Step 3: Run for 14 to 30 days. Fourteen days is the minimum for meaningful statistical signal on most fashion store traffic volumes. Thirty days controls for day-of-week variation and short-term novelty effects.
Step 4: Measure the right metrics. The metrics that matter:
Conversion rate (orders / sessions) for both groups
Average order value for both groups
Revenue per visitor for both groups
Return rate for both groups, measured 30 days after delivery to allow for return behavior
Revenue per visitor is the most useful single number, it combines conversion rate and AOV into one metric and controls for traffic volume differences.
Step 5: Calculate statistical significance. A result is meaningful when the difference between groups is large enough that it's unlikely to be random variation. A simple chi-square test on conversion rate, or a t-test on revenue per visitor, tells you whether the result would hold up if you ran the test again. Your vendor should be able to do this calculation; if they can't, that's a signal about how rigorously they measure their own performance.
What to ask your vendor before you sign
Most vendors will not proactively offer holdout data. You have to ask for it. Here are the questions that separate honest operators from ones who are hoping you won't dig.
"Can you show me holdout-based lift data from a live merchant deployment that's comparable to our store?" Not aggregate data. Not data from their best-performing merchant. A merchant comparable in volume, category, and price point to yours. If they don't have this or won't share it, their lift claims are based on attributed revenue, not incremental lift.
"How do you handle the attribution window?" Attribution windows are where vendor measurement tends to get creative. An AI interaction that occurs in a session that converts in the next 30 days can be credited to the AI, even if the AI interaction had nothing to do with the final purchase decision. Ask specifically: what counts as an "influenced" session, and over what time window?
"What is the average AI engagement rate across your merchant base?" If most shoppers aren't using the tool, the revenue impact can only be as large as the engaged fraction. A tool that's used by 4% of shoppers and converts them at 3x baseline is producing a net store-level lift of much less than 3x. The math matters.
"Do you separate novelty effect from steady-state performance?" New tools often see elevated engagement in the first weeks as shoppers explore something unfamiliar. This novelty effect inflates early metrics and fades. Honest vendors show 60 to 90 day data, not just the first two weeks.
"What happens to your numbers when you run a holdout?" Some vendors have run holdouts. Some haven't. Ask directly. A vendor who has run holdout tests and whose numbers held up will tell you. A vendor who hasn't run holdouts, or whose numbers dropped when they did, will change the subject.
The three metrics that actually matter for fashion AI
1. Revenue per visitor (RPV) — the composite metric
Revenue per visitor combines conversion rate and AOV into a single number and is the cleanest measure of what an AI shopping assistant is worth to your store. It controls for session volume and doesn't require you to separately interpret two numbers that may move in different directions.
A lift in RPV of 15% means every visitor to your store is worth 15% more in revenue when the AI tool is present. Applied to your monthly unique visitor count and average RPV, that's your monthly revenue increment.
Measure RPV for the AI-exposed group and the holdout group separately. The difference is your lift.
2. Return rate — the deferred metric
Return rate impact takes longer to measure because it depends on shopper behavior after delivery, not at the point of purchase. You need 30 to 45 days of post-delivery observation to capture the bulk of return behavior for fashion apparel.
Measure return rate per order for the AI-exposed group and the holdout group. Track by order date, not session date.
A genuine reduction in return rate means the AI tool isn't just converting more shoppers, it's converting them to purchases they keep. That's a margin story, not just a revenue story.
3. Repeat purchase rate — the long-term metric
The taste model behind a well-built AI shopping assistant is supposed to compound. A shopper who has used the system three or four times has given it enough signal to make meaningfully better recommendations than on visit one.
Measure repeat purchase rate at 60 and 90 days for shoppers who have had multiple AI-assisted sessions versus shoppers who haven't. If the taste model is working, you should see a meaningful difference in repeat purchase behavior by 90 days.
This is the metric most vendors don't talk about because it takes time and requires a long-term measurement commitment. It's also the metric where the compounding value of a system that actually learns shopper taste shows up most clearly.
What good numbers look like
To calibrate expectations: in well-run deployments with a proper holdout methodology, the typical AI shopping assistant shows:
15 to 35% lift in revenue per visitor on engaged sessions
10 to 25% reduction in return rate for shoppers with multiple AI interactions
20 to 40% higher 90-day repeat purchase rate for taste-model users versus non-users
These numbers are for tools that are actually working, engaged by a meaningful percentage of shoppers, with a genuine taste model, measured against a clean holdout. The "3x conversion" numbers often cited represent the lift on the specific slice of shoppers who actively engaged with the tool, not store-level lift. The distinction is important.
A tool that converts 3x on the 5% of shoppers who use it is producing a 10 to 15% store-level lift, which is real and valuable, but not the same as a 3x lift on your whole store.
Red flags in vendor measurement
These are the claims and configurations that should prompt more questions:
"AI-influenced revenue" usually means revenue from any session that included an AI interaction, regardless of whether the interaction was causal. Ask for incremental lift data.
Conversion rate presented without session comparison to holdout: a conversion rate on AI sessions is meaningless without comparison to a control group of similar shoppers who didn't have access to AI.
Attribution windows over 7 days: an AI interaction that's credited with a purchase that happened 28 days later is not a conversion; it's a coincidence dressed as an attribution.
"Our merchants average X% lift" — ask whether this is across all merchants or the ones they highlight. Survivor bias in vendor case studies is real; the merchants with negative results don't tend to appear in marketing materials.
No holdout methodology offered: if a vendor has been operating for more than a year and has never run a controlled experiment on their own impact, that's a signal about how confident they are in what a rigorous test would show.
The measurement posture that protects you
The simplest rule: don't let your vendor measure their own impact with their own methodology, reported through their own dashboard, without a controlled comparison group.
The AI shopping assistant vendor's incentive is to show you a large number. Your incentive is to know the true number. These are not the same. A vendor who offers holdout-based measurement is either confident in their results or honest enough to show you what they find regardless. Both of those are good signs.
Before deploying any AI shopping assistant at meaningful budget, agree in writing on the measurement methodology. Specify the holdout percentage, the measurement period, the primary metric, and who runs the analysis. Build the test before you build the expectation.
Elara's holdout-based lift report is built into every deployment. After 14 days, you receive a clean comparison between the exposed group and the holdout, conversion, AOV, and revenue, side by side. Your actual numbers, not ours.
