New research suggests that websites do not need a content agreement with OpenAI to appear in the index used by ChatGPT Search.
French SEO consultancy Resoneo analysed 1,249 ChatGPT responses captured during July and found that hundreds of websites without licensing agreements appeared to be served through OpenAI’s own search index in the same way as sites that have formal partnerships.
The findings also support a correction published earlier by SEO consultant Suganthan Mohanadasan, who initially believed OpenAI’s internal index mainly contained established publishers before revising his conclusion after further testing.
Small Sites Appear Alongside OpenAI Partners
Resoneo examined the different systems used to retrieve web results in ChatGPT and identified one known as “labrador”, which is believed to be OpenAI’s own search index.
When the consultancy compared pages retrieved through this system, it found no obvious difference between websites with OpenAI content agreements and those without them.
Pages from both groups appeared to have similar formats, snippet lengths and freshness.
Resoneo describes the index as being supplemented by press feeds and open scientific archives, allowing OpenAI to access some content directly rather than relying entirely on external search providers.
Its free-account data showed that questions with established answers, local businesses and products were frequently answered using this in-house index.
News searches were more evenly split between OpenAI’s index and results obtained through Google.
The situation was different for paid accounts using thinking mode. From 16,407 search results analysed by Resoneo, roughly 75% came through Google while about 24% were retrieved from OpenAI’s own index.
Earlier Findings Were Revised
The latest research follows an earlier investigation by Suganthan Mohanadasan into how ChatGPT selects its sources.
Mohanadasan initially suggested that OpenAI’s internal index appeared to operate as a type of approved list containing established publishers such as Reuters, The Guardian, the Wall Street Journal and Wikipedia.
However, he later revised that conclusion after receiving evidence from users showing that smaller publishers were also being served through the same system.
Further testing showed that publishers of different sizes could appear through the same retrieval pipeline.
Mohanadasan acknowledged that his original interpretation had gone too far and explained that his initial conclusion was based largely on data from a single account.
Resoneo’s research looked at a broader range of accounts and locations, including free and paid users and logged-out sessions. The researchers also replayed the same queries across different account types to compare the results.
According to Resoneo, OpenAI stopped exposing the retrieval-system label in its network traffic around 21 July. This means future investigations may have a harder time identifying exactly which system supplied individual results.
What ChatGPT Stores From A Web Page
The research also looked at how OpenAI’s index represents individual web pages.
Resoneo reviewed 534 pages cited by ChatGPT and compared them with the snippets stored in OpenAI’s index.
Of the 463 pages that contained an H1 heading, 387 had snippets that included the H1. That works out at around 83.6%.
The researchers found that snippets were generally cut off at around 200 characters. In many cases, the text appeared to come from the beginning of the page rather than the meta description.
This is different from Google’s approach, where the search engine can frequently use the meta description when creating a search snippet.
The median H1 length in the sample was 51 characters, leaving around 150 characters for the rest of the snippet.
Other elements displayed near the top of a page could also use up this limited space. For example, Resoneo found that:
- Around 29% of pages displayed a section label before the H1.
- About 11% included a publication date.
- Roughly 9% included the alt text of the first image in the snippet.
- Around one in seven pages had no H1 heading.
Where there was no H1, the snippet generally began with another heading or piece of content provided by the website template.
Why This Matters For Smaller Websites
The findings could be encouraging for smaller publishers and businesses that do not have agreements with OpenAI.
A website does not appear to need a formal content deal simply to be included in the index used to answer many free ChatGPT searches.
This challenges the idea that OpenAI’s own index is primarily reserved for major publishers or commercial partners.
However, being included in the index does not necessarily mean a website will receive more citations.
The research examined how pages were stored and retrieved, rather than whether having an OpenAI agreement increases the likelihood of being cited.
That distinction is important for businesses looking at AI search visibility.
Content Deals May Still Have Other Benefits
The findings do not mean that OpenAI’s publisher agreements have no value.
Content partnerships may provide OpenAI with information through direct feeds rather than conventional crawling, which could affect how quickly or reliably certain content reaches its systems.
There may also be other benefits associated with individual agreements that are not visible from public search results.
However, the research suggests that simply appearing in ChatGPT’s free search results is not necessarily one of the main reasons a publisher would need such an agreement.
Websites without a deal are already appearing through OpenAI’s internal retrieval system.
What Website Owners Should Take From This
For website owners and SEO professionals, the findings provide some useful context around AI search visibility.
Smaller websites should not assume that they are excluded from ChatGPT simply because they do not have a relationship with OpenAI.
Instead, it may be more useful to focus on making content easy for AI systems to understand and retrieve.
The research also highlights the importance of the content near the beginning of a page. If OpenAI’s index is storing a relatively short extract, headings and the opening section of an article could have a significant influence on what information is available to the system.
This does not prove that changing an H1 or opening paragraph will increase the chance of being cited. Resoneo did not test whether changes to page structure directly influence citation rates.
Nevertheless, it offers an interesting insight into how AI search systems may process website content.
AI Search Visibility Is Still Developing
The way ChatGPT finds and selects web content is continuing to evolve.
OpenAI’s crawler documentation does not currently provide a complete explanation of how its internal search index works or how publisher agreements affect retrieval and citations.
As a result, SEO professionals should be cautious about drawing broad conclusions from individual accounts or limited datasets.
The latest Resoneo research provides stronger evidence that OpenAI’s search index is not simply a closed system reserved for large publishers. Websites of different sizes can appear through the same retrieval process, even without a formal content deal.
For smaller websites, that means the opportunity to appear in AI-generated search results may be more open than previously thought.
More Digital Marketing BLOGS here:
Local SEO 2024 – How To Get More Local Business Calls
3 Strategies To Grow Your Business
Is Google Effective for Lead Generation?
How To Get More Customers On Facebook Without Spending Money
How Do I Get Clients Fast On Facebook?
How Do You Use Retargeting In Marketing?
How To Get Clients From Facebook Groups