Back to blog

Blogo Team

AI Search Optimization for B2B: What Is Known and a Page Checklist

What Google and OpenAI have published about AI search sources, what nobody outside them knows, and a pass/fail checklist for a B2B supplier site.

Google and OpenAI have each published what a page needs before their AI search features can use it, and the list is short: the page must be crawlable, indexed, and allowed to be shown. Neither company has published how one eligible page gets chosen over another. So the work you can verify is eligibility and substance, and no checklist, including the one below, can promise that an AI answer will cite your site.

What Google has published

Google's page on AI features and your website covers AI Overviews and AI Mode. Four statements in it matter to a supplier site:

  • No separate optimization. It says there are "no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary."
  • Indexed, with a snippet. To be shown as a supporting link, a page must be indexed and eligible to appear in Google Search with a snippet.
  • Eligible is not the same as shown. The same page says that meeting every requirement does not mean Google will crawl, index or serve the content.
  • Several searches per question. Both features may use what Google calls "query fan-out": running multiple related searches across subtopics to build one response.

Google's guide to optimizing for generative AI search adds what you can skip. It says you do not need to create new machine-readable files, AI text files, markup or Markdown to appear in Google Search, and it names llms.txt as an example. It also says there is no requirement to break content into tiny pieces. What it recommends is "non-commodity" content: material that does not repeat what others have already said or what a generative model could easily produce.

For measurement, the AI features page says traffic from these features is counted in the Search Console Performance report under the Web search type, and the generative AI guide points to a Generative AI performance report.

What OpenAI has published

OpenAI's crawler documentation describes separate user agents with separate jobs:

User agentWhat OpenAI says it is forWhat that means for you
OAI-SearchBotSurfacing websites in ChatGPT's search featuresSites opted out of it are not shown in ChatGPT search answers, though they can still appear as navigational links
GPTBotCrawling content that may be used to train its modelsBlocking it is a training decision, not a search decision
ChatGPT-UserVisits triggered by a user's action in ChatGPTOpenAI says it is not used to decide whether content appears in search

The same page recommends allowing OAI-SearchBot in robots.txt if you want your site to appear in search results, and says a robots.txt change can take about 24 hours to take effect for search.

OpenAI's help article on ChatGPT search describes answers that come with links to the web sources used. That is the citation you are hoping for, and the article is a description of the feature for users, not a guide for publishers.

What is not known

Be suspicious of anyone who claims certainty on these points, because neither company has published them:

  • How a source is picked among eligible pages. The documents above describe access and eligibility. They do not describe selection.
  • Whether a format wins. There is no published rule that tables, FAQ blocks, a particular word count or a particular heading pattern leads to citation. Google's guide explicitly says chunking is not required.
  • Whether being cited brings enquiries. A link in an answer is not a visit, and a visit is not an RFQ. You have to measure this on your own site.
  • How stable any of it is. Nothing published says an answer will cite the same sources next week. A single screenshot of your company in an answer is an observation, not a ranking.

We asked an OpenAI model with web search 100 questions that B2B marketers and export manufacturers ask and recorded which pages it cited. The write-up is in our ChatGPT citation study. Treat it the same way: one sample, on one set of dates.

What you can do: a pass/fail checklist

Everything here is either required by the documentation above or is ordinary good practice for a page a buyer would want to read. None of it is a lever that forces a citation.

CheckHow to test itPass when
Search crawlers are allowedOpen yoursite.com/robots.txt and read the rules for Googlebot and OAI-SearchBotNeither is disallowed from the pages you want found
The page is indexedInspect the URL in Search ConsoleThe result says the URL is on Google
A snippet is allowedView the page source and search for nosnippet and max-snippetNeither restricts the page, unless you chose that on purpose
The canonical is rightIn the inspection result, compare the user-declared and Google-selected canonicalBoth show the URL you want people to land on
The main content is in the HTMLUse the live test in URL Inspection and view the crawled pageSpecifications and tables appear as text, not only inside images or PDFs
One buyer task per pageRead the title and first paragraphYou can say in one sentence which question the page answers
Facts are stated plainlyRead each number on the pageEach has a unit, a condition and a source or a named owner
It adds somethingCompare it with the top results for the same questionIt contains at least one thing from your own production that they lack

Three of these need more explanation.

A page per buyer task. Because one question can fan out into several searches, a page that answers one narrow task completely is easier to match than a page that touches ten topics lightly. "Which closure liner for an oil-based serum" is a task. "Everything about closures" is a category. Write the first kind as articles and keep the second as navigation.

Facts stated plainly, with sources. "Our caps are highly durable" gives a reader, human or machine, nothing to use. "Example: torque retention measured after 24 hours at room temperature, method described below" does. Where the fact comes from a standard or a regulator, link to that body's own page. Where it comes from your own testing, say so and say how it was measured. Blogo applies this rule when it drafts: each factual statement is checked against a source, and a statement nothing supports is removed or rewritten.

Canonical correct. Product catalogues often expose the same content at several URLs through filters, print views or tracking parameters. Google's page on consolidating duplicate URLs lists redirects and rel="canonical" as strong signals and sitemap inclusion as a weak one, and tells you to link internally to the canonical URL. If Google has selected a different canonical from the one you declared, fix that before you work on anything else on the page.

Is a blog still worth writing?

It depends on what the blog says. Google's guide draws the line itself: content a generative model could easily produce is the kind to avoid. A post that explains what a ball valve is competes with the model's own answer. A post that states the tolerance you hold on a part, how you test it, and what fails when a buyer specifies the wrong material contains information the model does not have unless it reads your page.

That is also the content a purchasing engineer wants before sending an enquiry, so the work is not wasted if AI search never sends you a visitor. Write for that reader, keep the pages eligible using the checklist, and measure enquiries. Do not measure screenshots.

Want Blogo to draft your next article?

Start with your website. You check every fact before anything goes to WordPress.