AI Search Visibility Using SEO Datasets
Publishing SEO datasets improved the one part of AI search visibility I can measure from my own server. AI crawlers and search engines fetched the data files within days of publishing, and the logs show them following the link from the page to the file. It has not earned a single Google click on a dataset page, and no server log can show a citation.
Both halves are true. The gap between them is this case study. It covers 35 days of server logs across 190 websites and four months of Search Console data, including the day catalog fetches more than doubled and the strangest query Google ever sent to one of my pages.
A case study from the Digital Karma Data Warehouse on which crawlers fetch published SEO datasets, how they find the files, what Google does with a dataset page, and what none of it proves.
What Counts as an SEO Dataset in This Case Study
An SEO dataset here is our own search and crawler data, published in three forms at once:
- A page a person can read, with the finding and a table.
- A JSON file anyone can download and reuse.
- Dataset structured data on the page that names the file, so a crawler reading the page knows the file exists.
AI Symantix published five of them in September 2026. All five come from the Digital Karma Data Warehouse, the private database where Search Console data and server logs for every site in the portfolio land each night.
| # | Dataset | Published | Page | File |
|---|---|---|---|---|
| 1 | CTR by Search Position | September 7, 2026 | /datasets/ctr-by-position/ | /ai/ctr-by-position.json |
| 2 | Search Visibility Index | September 7, 2026 | /visibility-index/ | /data/visibility-index/index.json |
| 3 | AI Bot Visibility on one site | September 11, 2026 | /ai-bot-visibility/ | /data/research/ai-bot-visibility.json |
| 4 | AI Bot Visibility across five websites | September 11, 2026 | /ai-bot-visibility-cohort/ | /data/research/ai-bot-visibility-cohort.json |
| 5 | Search Visibility to Inquiries baseline | September 12, 2026 | /search-visibility-to-inquiries/ | /data/research/search-conversion-baseline.json |
Then the whole portfolio joined in. On September 23 every page on all 190 websites started carrying a DataCatalog node and Dataset nodes in its JSON-LD, each pointing at that site's /ai/catalog.json. That is the Digital Karma Federation v8.2 rule, published at DigitalKarmaWeb.com. It matters here because it handed me a before and an after.
Three sister sites do related work. DatasetSEO.com shows dataset-backed website systems with their measurements. DatasetsMaker.com helps you build your own. AI to AI Datasets catalogs datasets built for AI systems to consume. They show up again below, because the logs did not stay inside one domain.
Crawlers Took the File the Page Pointed To
The first thing I wanted to know was simple. When a crawler reads the page, does it take the file?
Server logs record a referer, the address a crawler says it came from. For the dataset files on AI Symantix, the referer was almost always the page that links the file.
| File fetched | Crawler | Fetches in 35 days | Where it came from |
|---|---|---|---|
| /data/lab/exposure-velocity.json | Googlebot | 18 | /ai-symantix-lab/ |
| /data/lab/exposure-velocity.json | GoogleOther and Applebot | 3 | /ai-symantix-lab/ |
| /data/research/search-conversion-baseline.json | PetalBot | 3 | /search-visibility-to-inquiries/ |
| /data/research/ai-bot-visibility-cohort.json | GPTBot and PetalBot | 3 | /ai-bot-visibility-cohort/ and /datasets/ |
| /data/research/ai-bot-visibility.json | GPTBot and PetalBot | 2 | /ai-bot-visibility/ |
| /ai/ctr-by-position.json | PetalBot | 2 | /datasets/ |
| /data/visibility-index/index.json and latest.json | GPTBot | 2 | the server's preview copy of /visibility-index/ |
| /data/research/search-conversion-baseline.json | GPTBot | 1 | www.signalarchitectgroup.com |
Googlebot fetched the Exposure Velocity file 18 times in 35 days, and every single time it came from the Lab page. GPTBot read the five-site cohort page, then took the cohort file. PetalBot, the search crawler from Huawei, did the same thing on six different files.
Nobody submitted anything anywhere. The page linked the file. The crawler took it.
Speed was not a problem either. The CTR by Search Position file went live on September 7 and was fetched the same day by a bot I could not name. GPTBot read the dataset page on September 10. Bingbot had the file by September 15 and came back for it six more times.
The last row in the table is the one I keep looking at. On September 12 the Search Visibility to Inquiries study was linked from Signal Architect Group, a sister site. That same day GPTBot fetched the baseline file on AI Symantix with signalarchitectgroup.com as its referer, and it fetched DatasetSEO's copy of the same study the same way. One link on a sister site moved an AI crawler across three domains in a day.
This is the same behavior RealSEOLife recorded on an IBM software site in August, when a GPTBot client took eleven data files in 92 seconds that were declared only in JSON-LD. That receipt is Eleven Machine-Readable Resources in 92 Seconds. Two sites, two months apart, same pattern.
Who Fetches SEO Datasets Across 190 Websites
One site is an anecdote. So I pulled the same question across every site in the portfolio for the 35 days of logs the warehouse keeps, September 6 through October 10, 2026.
The dataset layer in these counts means five kinds of file:
- The catalog at /ai/catalog.json.
- The manifest at /ai/manifest.json.
- The llms.txt and llm.txt guidance files.
- The federation files for health and relationships.
- Every public JSON data file under /ai/, /data/ and /content/.
Only successful GET requests count.
| Who fetched the dataset layer | Requests | Sites where the catalog was fetched |
|---|---|---|
| Bots I could not name | 11,969 | 176 |
| People | 3,390 | 107 |
| SEO tools, monitors and other named bots | 3,252 | 113 |
| Search engine crawlers | 1,463 | 128 |
| AI training crawlers | 1,384 | 143 |
| AI search crawlers | 244 | 57 |
| AI tools fetching for a person | 11 | 1 |
That is about 21,700 requests to the dataset layer. The same 190 sites served 7.09 million content page requests in the same window. The dataset layer is three tenths of one percent of the crawl.
Two things in that table surprised me.
People fetched llms.txt and llm.txt 1,853 times from 1,469 different addresses. I wrote those files for AI systems. Humans are reading them.
And the biggest group has no name. Nearly 12,000 requests came from bots the warehouse classifies as bots but cannot match to any known crawler. They hit the catalog on 176 of 190 sites. I do not know who they are, so I am not going to pretend they are AI search.
OAI-SearchBot Spends One Visit in 25 on the Dataset Layer
Raw counts favor the crawlers that crawl the most. The better question is what share of each crawler's visits went to the dataset layer instead of ordinary pages.
| Crawler | Owner | Job | Dataset-layer fetches | Page fetches | Share | Sites with a dataset fetch |
|---|---|---|---|---|---|---|
| OAI-SearchBot | OpenAI | AI search | 222 | 5,414 | 3.9% | 62 |
| Applebot | Apple | Search | 141 | 9,020 | 1.5% | 52 |
| Bingbot | Microsoft | Search | 294 | 20,051 | 1.4% | 67 |
| CCBot | Common Crawl | AI training | 25 | 2,098 | 1.2% | 10 |
| PetalBot | Huawei | Search | 439 | 43,125 | 1.0% | 34 |
| Googlebot | Search | 223 | 59,219 | 0.37% | 78 | |
| PerplexityBot | Perplexity | AI search | 20 | 7,564 | 0.26% | 10 |
| GoogleOther | Search | 172 | 85,170 | 0.20% | 83 | |
| ChatGPT-User | OpenAI | Fetching for a person | 10 | 8,495 | 0.12% | 4 |
| GPTBot | OpenAI | AI training | 942 | 2,188,904 | 0.04% | 155 |
| meta-externalagent | Meta | AI training | 211 | 809,146 | 0.03% | 41 |
| ClaudeBot | Anthropic | AI training | 63 | 292,151 | 0.02% | 37 |
| Amazonbot | Amazon | AI training | 105 | 579,240 | 0.02% | 22 |
Read the top row and the GPTBot row together. OAI-SearchBot, the crawler OpenAI uses for search answers, spends about one visit in 25 on the dataset layer. GPTBot, the crawler the same company uses for training, spends about one visit in 2,300. Same company. Two very different appetites.
The search crawlers from Apple, Microsoft and Huawei sit near the top too. The training crawlers from OpenAI, Meta, Anthropic and Amazon sit at the bottom. The pattern is clear enough to say out loud: crawlers that answer questions want the catalog. Crawlers that collect text want pages.
Googlebot sits in the middle, and that fits the referer table. Google takes a data file when a page it already trusts points at one. It does not go looking for catalogs.
Catalog Fetches Doubled After Every Page Declared Its Datasets
September 23 split the log window almost in half. Before that date, most sites declared their datasets in one place, the catalog itself. After it, every page on every site carried the DataCatalog and Dataset nodes that point at the catalog.
So I counted how many different sites each AI crawler fetched the catalog on, in the 17 days before and the 18 days after. The BellyUp city sites are left out of this table for a reason explained two sections down.
| AI crawler | Sites with a catalog fetch, September 6 to 22 | Sites with a catalog fetch, September 23 to October 10 | Catalog fetches before | Catalog fetches after |
|---|---|---|---|---|
| GPTBot | 54 | 87 | 105 | 127 |
| OAI-SearchBot | 11 | 43 | 14 | 44 |
| meta-externalagent | 3 | 19 | 10 | 28 |
| ClaudeBot | 5 | 6 | 5 | 7 |
OAI-SearchBot went from 11 sites to 43. Meta's crawler went from 3 to 19. Across the whole portfolio, GPTBot's busiest catalog day before the change touched 11 sites. On September 24 it touched 28. On September 25 it touched 30.
I want this to be proof. It is not. The after window is one day longer, crawlers run on schedules I cannot see, and GPTBot crawled far more of everything in October. Counting sites instead of requests takes most of that away, and the jump shows up in crawlers from three separate companies in the same two days. I am calling it a strong hint. The next monthly reading will say whether it held.
Google Lists the Dataset Page as a Page, and the Queries Are Odd
Search Console tells a different story, because it measures a different thing. Google shows a dataset page for the words on it, like any other page. Here is every AI Symantix dataset page Google has shown since the pages went live.
| Page | First shown | Impressions | Clicks | Average position |
|---|---|---|---|---|
| /ai-symantix-lab/ | August 27 | 66 | 0 | 4.7 |
| /datasets/ctr-by-position/ | September 14 | 54 | 0 | 58.4 |
| /datasets/ | August 10 | 44 | 0 | 38.2 |
| /ai-bot-visibility/ | September 12 | 28 | 0 | 52.1 |
| /visibility-index/ | September 8 | 23 | 0 | 19.0 |
| /search-visibility-to-inquiries/ | September 12 | 7 | 0 | 13.9 |
| /ai-bot-visibility-cohort/ | September 12 | 4 | 0 | 27.5 |
Zero clicks on every row. The Lab page sits at position 4.7 because the only query it ranks for is the brand name.
Then there is the query that landed on the CTR dataset page 27 times, at position 52:
top queries,clicks,impressions,ctr,position
That is the header row of a Search Console export, pasted into Google with the commas still in it. I do not know whether a person did that or a script did. It reads like a script. Either way Google decided a page that publishes click-through rate by position might answer it, and it was not wrong.
The plainer queries were "ctr by position" at position 66 and "organic ctr by position" at position 95. Right subject, wrong depth.
The phrase in this case study's title is a real search too. Over on DatasetsMaker, the SEO Keyword and Ranking Dataset page has 797 impressions and one click since July. The query "seo datasets" gave it 171 impressions at position 62. The query "seo dataset" gave it 130 at position 74. People look for SEO datasets. Google shows ours at depths where nobody clicks.
One more thing Search Console will not do is count any of this in its Dataset report. That report only sees pages carrying @type: Dataset in their structured data, which is the whole point of the RealSEOLife piece Dataset SEO: The Schema Label Google Is Actually Looking For. To Google, the file is not the dataset. The label is.
What the SEO Datasets Did for AI Search Visibility on AI Symantix
The datasets were one part of a bigger push on this site. The checker, the glossary cluster and the page consolidation all landed in the same weeks, so I cannot hand the datasets their own share of what moved. I can show what moved.
| Week of | Impressions | Clicks | Weighted position | Queries |
|---|---|---|---|---|
| June 16 | 236 | 0 | 88.2 | 25 |
| July 13 | 1,142 | 0 | 81.9 | 71 |
| August 3 | 1,488 | 1 | 60.2 | 111 |
| August 31 | 1,498 | 1 | 56.3 | 101 |
| September 7 | 1,600 | 0 | 41.1 | 96 |
| September 14 | 1,568 | 2 | 34.5 | 90 |
| September 28 | 976 | 2 | 32.7 | 92 |
| October 5 (four days) | 904 | 1 | 30.4 | 103 |
Weighted position went from 88 to 30 in four months. Impressions climbed to 1,600 a week and then fell back to about 900 as the deep impressions nobody sees dropped away. Clicks stayed between zero and two a week the whole time.
The AI search visibility queries moved further. In the week of June 16 they averaged position 89.9. In the first week of October they averaged 20.9, on fewer impressions, because Google stopped showing the pages at position 80 for things they do not answer.
Here is where publishing the dataset paid off in a way I did not plan. The CTR by Search Position dataset says a page at positions 21 to 50 earns a click on 0.15 percent of impressions. At positions 11 to 20 it is 0.49 percent. On the top three spots it is 5 to 17 percent. At 900 impressions a week and position 30, the curve predicts one or two clicks. That is exactly what the site gets.
The dataset I published to explain search visibility explained my own.
And the five-site cohort study had already warned me not to read crawler growth as a forecast. AI crawler visits to this site went from 362 in June to 1,754 in September. Clicks did not follow, because crawler visits are retrieval, and retrieval is only the first of three parts.
People did show up. Between 14 and 20 different visitors opened each dataset page in the 35 days, almost all of them with no referer. That usually means a link in an app, an email or an AI chat. The logs cannot say which.
The Biggest Number in the Warehouse This Month Is a Trap
AI crawlers made 763,949 requests across the portfolio in September. In the first 11 days of October they made 3,571,310.
I could have opened with that number. It would have been a lie with a chart.
The top twelve sites in October are all BellyUp city restaurant finders. The bars lens of BellyUp Tampa alone took 310,575 AI crawler requests in 11 days. GPTBot made 1.8 million of the month's requests. Meta's crawler made 783,000 and Amazonbot made 645,000. That is crawlers looping through every filter combination a finder can produce. RealSEOLife documented the first one of these in a 43,000-request crawl trap, and the pattern is back at a bigger scale.
So the before and after table above leaves the BellyUp sites out, and no number in this case study depends on raw AI crawler volume. A request proves a request. It does not prove interest.
What SEO Datasets Cannot Show
AI search visibility has three parts: retrieved, cited and accurately represented. Everything above lives in the first part. A fetch in the log proves a crawler asked for a file and got it. It does not prove the file was kept, used in an answer or described correctly.
The closest thing to a citation signal in the logs is an AI tool fetching a file while it answers a person. ChatGPT-User and Claude-User did that 11 times across four sites in 35 days. That is not nothing, and it is not a citation either. An AI citation only shows up when you ask the questions and read the answers, which is how the guide to checking AI search visibility says to test it.
The glossary entry for AI findability covers the retrieval side in more depth. Can AI Bot Data Measure AI Search Visibility? covers why crawler counts stop where they stop.
How to Use SEO Datasets for AI Search Visibility
Five things the data supports, in the order I would do them:
- Publish each dataset three ways at once: a readable page, a downloadable file, and structured data that names the file. Each form reaches a different reader.
- Link the file from the page in plain HTML. The referer records show that the link is how Googlebot, GPTBot and PetalBot found every file.
- Put the catalog on every page, not only the homepage. The number of sites where AI crawlers fetched the catalog more than doubled after we did that.
- Measure with server logs and Search Console side by side. Each one is blind to what the other sees.
- Expect retrieval, not clicks. Clicks still follow position, and a CTR curve will tell you when to expect them.
Method and Limits
- Server logs cover September 6 through October 10, 2026, the 35 days the warehouse keeps. Search Console data for AI Symantix runs from June 16 through October 8, 2026.
- The portfolio means the 190 websites in the Digital Karma registry that served content in the window. Sites outside the portfolio are excluded.
- Only GET requests with a successful status count. Crawlers are named from the warehouse's fingerprint table and grouped by the job their owner publishes for them: search, AI search, AI training or fetching for a person.
- Bots I could not name are requests the warehouse classifies as a bot without matching a known crawler. Their identity is unknown and they are reported separately.
- Weighted position is position weighted by impressions. Weekly rows start on Monday. The October 5 row covers four days.
- Human counts may include our own visits. A same-day match between a link and a fetch is a sequence, not proof of cause.
- Every table on this page is a snapshot as of October 11, 2026. The public data file is /data/research/seo-datasets-ai-search-visibility.json. A parallel case study with the warehouse receipt is on RealSEOLife.com.
Check whether AI can retrieve your page
Paste in a URL and get a score, every pass and fail, and a fix for each failure. It is free and takes about a minute.
Open the AI Search Visibility CheckerFrequently Asked Questions
Do SEO datasets improve AI search visibility?
They improve the retrieval part. Within days of publishing, search engines and AI crawlers fetched the files and followed the page links to them. The datasets did not earn clicks on their pages and cannot show a citation.
Which crawlers fetch published SEO datasets?
In 35 days across 190 websites, GPTBot fetched the dataset layer on 155 sites and OAI-SearchBot on 62. Bingbot, Googlebot and PetalBot fetched it too. OAI-SearchBot spent about 4 percent of its visits on the dataset layer, the highest share of any named crawler.
Does Google index a dataset page?
Yes, as an ordinary page. The AI Symantix CTR dataset page has 54 impressions at an average position of 58 and no clicks. The Search Console Dataset report is separate and only counts pages that carry Dataset structured data.
Did AI crawlers fetch the catalog more after Dataset structured data went on every page?
Yes. Leaving out the BellyUp sites, OAI-SearchBot fetched the catalog on 11 sites in the 17 days before and 43 sites in the 18 days after. GPTBot went from 54 sites to 87 and the Meta crawler from 3 to 19. It is a strong hint rather than proof.
Can a dataset fetch prove an AI citation?
No. A fetch proves a crawler asked for the file and got it. A citation only shows up when you ask AI tools the question and read the answer.