All articles

We stopped guessing whether AI visibility work was doing anything. Here's the tool that tells us.

Before and after, on the same set of questions a real shopper would ask. We ran it on two client catalogs this month and watched visibility roughly double. Here's what the tool actually measures.

Veristyle AI helps fashion brands become part of those answers, by helping AI understand your products, your positioning, and your brand, because discovery is changing.

For most of this year, when a client asked us “did the optimization actually work,” the honest answer was some version of “give it a few weeks and check your traffic.” That’s a bad answer. It’s slow, it’s indirect, and it doesn’t tell you which specific product or category still isn’t showing up. So we built something that answers the question directly, before you wait on anything: a before-and-after test that runs real AI shopping questions against a catalog, twice, and shows you exactly what changed.

We’re calling it the Visibility Lab, and we’ve now run it on a couple of client catalogs this month. The results are the clearest evidence we’ve had yet that this work is measurable, not just plausible.

What it actually does

The idea is simple, even if the mechanics underneath aren’t. We take a set of realistic shopper personas, the kind of person who’d actually buy from a given brand, and generate the questions they’d plausibly ask an AI assistant. Not generic queries like “best coat,” but the specific, occasion-driven kind real shoppers use: what to wear to a fall wedding as a guest, which tote actually holds up for daily commuting, what a first date outfit looks like for someone who runs cold.

We run that full question set against the catalog as it exists today, and log whether each product gets retrieved at all, and if it does, whether the assistant actually recommends it. Then we do the optimization work, the product data enrichment we always talk about, and run the identical question set again. Same questions, same personas, same model. The only thing that changed is the catalog.

What we found

On one catalog we tested, a contemporary ready-to-wear label, the products started out being recommended in 4 of 28 realistic shopping queries. After enrichment, that number moved to 8. Doubling isn’t a rounding error. It’s the difference between an assistant recommending your product for one out of every seven relevant questions versus one out of every three and a half.

That gap matters more than the raw number, honestly. A brand that only gets recommended in 4 of 28 scenarios isn’t invisible, which is what makes this easy to miss. It’s just badly under-visible in a way that never shows up as a clean, obvious problem. Nobody’s product page is broken. Nothing throws an error. The catalog just quietly doesn’t answer enough of the questions a real shopper is asking, and the only way to know that was true was to actually ask.

Retrieved but not recommended is its own problem

One thing the early runs made obvious: “not showing up” and “showing up but getting passed over” are two different failures, and they need different fixes. A product that never gets retrieved usually has a data problem, thin descriptions, missing attributes, nothing for the model to match against the question. A product that gets retrieved but doesn’t get recommended usually has a persuasion problem. The model found it, understood roughly what it was, and still picked something else.

We’re building the Lab out to separate those two cases explicitly in the results, because right now we can tell a client “here’s what moved,” but we want to be able to tell them “here’s why the ones that didn’t move, didn’t,” down to the specific reason. That’s the next release.

Why we built this instead of just trusting the process

We spent a long time telling brands to trust that better product data leads to better AI visibility, and it does, but “trust the process” is a weak thing to ask a marketing team to do with their budget. A dashboard that shows before-and-after on the exact questions their customers are asking is a much easier thing to trust, because there’s nothing to take our word for. You can read the questions yourself and check whether the after column looks right.

It’s also just a better way to prioritize the work. If a test run shows a brand’s activewear category recommending well but its outerwear falling flat, that tells you exactly where to spend the next two weeks, instead of enriching the whole catalog evenly and hoping.

Try it on your own catalog

If you’re already a Veristyle client, ask your account contact to run a Visibility Lab test before your next optimization push, it’s the fastest way to see where the gaps actually sit before you spend time closing them. If you’re not yet a client and want to see what your own products look like on this test, a free AI visibility audit is the place to start, or book a demo and we’ll walk your catalog through it live.

Related reading

See Veristyle on your catalog

Book a demoGet it on Shopify