
Fix your schema, get more press, tidy your About page. None of it helps if AI has filed you under the wrong category. Change one word in the question and lululemon goes from being recommended 90% of the time to never.
Large language models do not decide whether your brand is good enough to recommend. They match the category word in someone's question against the category they have already filed you under, and that filing is built almost entirely from what other people have published about you. If the language your buyers use to describe what you do differs from the language your third-party coverage uses, you are invisible for exactly the questions that matter most.
That is a positioning problem wearing a technical costume. It belongs to marketing, not to whoever owns the website.
This guide covers what the research actually found, why the mechanism works the way it does, how to audit your own brand in about twenty minutes, and what to do when you find a gap. It also covers the limits of the evidence, because the honest version is more useful to you than the confident one.
Two things, working together.
The first is how you are described in structured, machine-readable sources. Google's Knowledge Graph carries a short description field for most recognised entities. Nike, New Balance and Reebok all carry the same one: footwear company. That description acts as an anchor. It tells the model, in the crudest possible terms, what kind of thing you are.
The second is the body of third-party content that has built up around your brand over time. Reviews, editorial comparisons, roundups, analyst pieces, best-of lists, comparison articles on other people's sites. This is where the detail lives. The anchor says which shelf you sit on; the corpus describes what you look like on it.
Maryanna Franco and Joao da Silva, who published the study this guide draws on, call the combination of the two category coding. The anchor drives recognition. The corpus drives recommendation. Most brands invest heavily in the first and assume the second follows. It does not.
Franco and da Silva tested twelve athletic apparel brands available in the UK across ChatGPT, Gemini, Perplexity, Claude and Google AI Overviews. Around 14,140 API runs over seven days. They then changed a single variable: the category word in the prompt. Athleisure in one condition, athletic footwear in the other.
The results moved in both directions at once.
New Balance was almost entirely absent from athleisure answers and almost universally present in footwear answers. lululemon was the exact reverse. Neither brand changed. Neither brand's authority changed. The models did not change. One word in the question changed, and the recommendation set turned over.
The interesting case is Nike, which held up in both framings at 77% and 90%. Nike's Knowledge Graph description is the same footwear company as New Balance's. What differs is the corpus. Nike has accumulated years of fashion and lifestyle coverage sitting alongside its performance coverage, so it registers as eligible in both categories rather than one. It built a second stream. New Balance did not.
Because they are answers to different questions.
Recognition asks: does the model know this brand exists and what it broadly is? Nike, New Balance and Reebok all pass that test perfectly. Their Knowledge Graph scores in the study ranged from around 666 to over 64,000, and all three were recognised without difficulty by every assistant tested.
Recommendation asks something narrower: when someone phrases a question this particular way, does this brand belong in the answer? To decide that, the model reaches for content that discusses brands using that phrasing. If your brand does not appear in that content, you are not a weaker candidate. You are not a candidate.
This is the part most brands get wrong. The instinct is to treat low AI visibility as a strength problem, so the response is more press, more authority, more coverage. But if the coverage keeps speaking the language of the category you are already coded into, you have made yourself more visible in the queries where you were already winning and no more visible in the queries where you were absent.
Strength is not the lever. Language is.
No, and it is worth understanding why not, because this is the shortcut everyone reaches for.
Editing the description changes the anchor. It does not change the corpus. If every article, review and roundup that mentions your brand talks about performance footwear, then relabelling yourself an apparel company gives the model a new anchor with nothing attached to it. The content still says what it always said, and the content is what gets retrieved when someone asks a question.
The same logic applies to your own website. Rewriting your homepage to use the new category language is necessary and worth doing, but your site is one source among thousands. The category is decided in aggregate, mostly by people who do not work for you.
That is uncomfortable, and it should be. It means the correction is slower and more expensive than a copy change.
You can run a usable version of this audit in about twenty minutes. It is deliberately crude. The point is to find the gap, not to produce a research paper.
Step one: write down five or six ways a buyer might phrase the category. Not your positioning statement. The words they would actually type. For a B2B software brand that might be customer onboarding software, user activation tools, product-led growth platforms, customer success software and in-app guidance tools. Some of these will feel like adjacent categories. Those are the ones worth testing.
Step two: turn each into a recommendation question. What are the best [category] for mid-market B2B companies? Keep the phrasing identical across all of them apart from the category words.
Step three: run each question across three assistants. ChatGPT, Claude and Perplexity is a reasonable spread. Run each one two or three times, because outputs vary between runs. Note whether you appear, and who appears alongside you.
Step four: build a simple grid. Category phrasings down the side, assistants across the top, present or absent in each cell. The pattern usually shows up immediately. You will appear consistently for one or two phrasings and be missing entirely from others.
Step five: for every phrasing where you are absent, ask three questions. Does any third-party content about your brand use that language? Are you covered in publications that write about that category? Do you appear in comparison pieces and roundups that use that phrasing?
If all three answers are no, you have found the gap and you know its shape.
One caveat on method. Assistants personalise and vary between runs, and a single absence proves very little. What you are looking for is a consistent pattern across repeated runs and multiple assistants, not a single missing mention.
The corrective lever the research points to is third-party content investment in the specific category framing your buyers use. In practice that means four things, in roughly this order.
Get the language right on your own surfaces first. Homepage, product pages, founder bios, LinkedIn descriptions, conference abstracts, podcast blurbs. Use the same category phrase verbatim everywhere and do not vary it for freshness. Consistency is the entire point. This will not fix the gap on its own, but every other tactic works better once your own definition is unambiguous.
Earn coverage in the publications that write about that category. Not more coverage generally. Coverage in the specific places a model would retrieve from when someone asks about that category. That usually means original data, a genuine point of view, or an expert who is worth quoting.
Get into comparison and roundup content that uses the phrasing. These pieces are disproportionately influential because they are structurally exactly what a recommendation question needs: a list of brands, in a named category, with reasons. If the roundups for your target category do not include you, that absence is doing real damage.
Sit alongside the brands that already define the space. Co-mentions matter. Being named in the same paragraph as the established players in a category is a strong signal that you belong in it.
Then re-run the audit in three months. This is slow work, and there is no version of it that pays off in a fortnight.
Less than the headline numbers invite you to.
The study is a controlled test with one variable changed, which is genuinely rare in this area and makes the finding more credible than most of what gets published about AI visibility. The symmetry of the results, moving roughly equally in both directions, is hard to explain as noise.
But it covers twelve brands, in one consumer vertical, in one market, over seven days. Consumer athletic apparel is a category with unusually heavy editorial and lifestyle coverage, which may make the corpus effect stronger there than it is in B2B software, where the third-party content layer is thinner and analyst reports carry more weight. Nobody has run the equivalent study on B2B categories yet, as far as we can tell.
There is also a wider disagreement worth knowing about. Research on whether structured data affects AI visibility has produced contradictory results, with some studies finding no relationship at all. The most defensible reading is that structured data is infrastructure for machine understanding rather than a route to citations, and that visibility follows from being understood rather than from being marked up.
So: treat the mechanism as well-argued and the magnitude as unproven outside the tested category. The audit is worth twenty minutes regardless, because it costs almost nothing and tells you something specific about your own brand rather than about athletic apparel.
It moves a chunk of AI visibility work out of the technical column and into yours.
The questions this raises are positioning questions. Which category do we say we are in? Which category do buyers say we are in? Which category do the analysts, the roundups and the comparison pieces put us in? If those three answers differ, that is not a search problem to be delegated. That is a positioning problem that has always existed and is now being priced by machines.
There is a version of this that is genuinely useful for anyone who has been trying to explain to a leadership team why consistent category language matters. The argument used to rest on human memory and mental availability, which is true but slow to demonstrate. Now there is a measurable version: change the word, watch the recommendation set turn over.
It also raises an uncomfortable question about adjacent categories. If your buyers are drifting towards new language, and your coverage is not drifting with them, the gap opens quietly. Nothing breaks. You simply stop appearing in the questions people are starting to ask. This is the kind of judgement call that separates the marketers described in What Is an AI-Native Marketer? (And How to Become One) from those still treating AI as a drafting tool.
What is category coding?
Category coding is the combination of two things that determine whether an AI assistant recommends a brand for a given question: the short structured description that anchors the brand to a category, and the body of third-party content that has accumulated around the brand in that category. The first drives recognition; the second drives recommendation.
Why does AI recommend my competitor and not me?
Most often because your competitor's third-party coverage uses the category language in the question and yours does not. It is usually a language mismatch rather than a judgement about quality or authority.
Does schema markup improve AI visibility?
The evidence is contested. Microsoft has confirmed that Copilot uses structured data to understand content, while independent studies have found no measurable effect on citations. The most defensible position is that structured data helps machines understand what you are, and visibility is a downstream effect of being understood accurately. It is not a route to citations on its own.
How often should I run a category audit?
Quarterly is sensible for most brands. More often than that and you are measuring run-to-run variance rather than real change, given how slowly third-party content accumulates.
Can I fix this by changing my website copy?
Not on its own. Your own surfaces should use consistent category language, and that is worth doing first because everything else works better afterwards. But the category is decided largely by content you do not own, so the correction requires earning coverage in the language your buyers use.
Does this apply to B2B as well as consumer brands?
The mechanism should, but the study tested consumer athletic apparel only. B2B categories have a thinner third-party content layer, which may mean individual pieces of coverage carry more weight, or may mean the effect is weaker. This has not been tested yet.
Run the audit. Five or six phrasings, three assistants, twenty minutes, one grid. It is the cheapest diagnostic in this whole area and most brands have never done it.
For the wider picture of where this sits alongside everything else AI is changing in B2B marketing, see AI in Marketing: A B2B SaaS Marketer's Field Guide.
If you want the detail on how to build the third-party content plan that closes a category gap, that is the kind of work we go through in depth inside SaaStrix, the AI-native marketing community for B2B marketers. Members get the audit template, the worked examples and a room of people running the same play on their own brands.
SaaStrix is where B2B marketers become agentic marketing leaders, and you can try it free for five days.
Source: Franco, M. and da Silva, J., The recognition-recommendation gap: Empirical evidence that category coding, not knowledge-graph strength, determines brand visibility in generative AI output. Published open access on Zenodo (DOI 10.5281/zenodo.20331344). Findings summarised in Search Engine Land, 21 July 2026.