AI visibility has become one of the biggest talking points in SEO. Over the past few years, businesses have focused heavily on making their websites easier for AI systems to crawl, read and reference.
If your brand is already appearing in AI-generated answers, earning citations or being included in AI Overviews, it can be tempting to think the hard work is largely done.
But being visible to AI is about more than simply getting mentioned.
An audit of 50 major websites found a significant gap between making content accessible to AI and making that content genuinely understandable and usable by AI systems. While many websites have improved their technical accessibility, very few have built the infrastructure needed to help AI accurately understand their content.
There is another issue too. Almost two-thirds of the websites examined appeared to have no clear strategy for deciding which AI crawlers should be allowed to access their content.
This suggests that AI search optimisation is moving into a new phase. Traditional SEO principles such as crawling and indexing still matter, but businesses now need to think about how AI interprets their websites and, increasingly, how AI agents can interact with them.
Three Layers of AI Visibility
AI visibility is often measured by how frequently a company, product or website appears in AI-generated responses. Mentions and citations are certainly useful, but they represent only one part of the picture.
For AI systems to make reliable use of a website, there are three separate stages to consider: finding the information, understanding it correctly and being able to take action with it.
1. Retrievability
The first question is simple: Can an AI system find and process your content?
This is the area most closely connected to traditional technical SEO.
A website needs to provide clean HTML, a sensible structure and accessible content so that AI crawlers can retrieve information without encountering unnecessary obstacles.
Technical improvements such as semantic HTML, accessibility labels, server-rendered content and appropriate crawler instructions can all contribute to this layer.
Retrievability is essentially the foundation. If AI cannot reliably access your content, it has little chance of understanding or using it properly.
2. Attribution and Meaning
The second layer goes beyond simply reading the information.
AI needs to understand what the information represents, who created it and which entity it relates to.
For example, a website might contain several prices for a product. AI needs to distinguish between the standard price, a discounted price and a price available only to members. Similarly, it needs to understand whether a page refers to a particular brand, company, person or organisation.
This becomes particularly important when names or terms have multiple meanings. A search for “Jaguar”, for example, could refer to the car manufacturer, an animal or a football club.
Structured data and other semantic signals can help AI establish these relationships rather than simply making an educated guess.
The benefit is twofold. Better interpretation can increase the chances of your content being included in relevant AI answers while also reducing the risk of AI systems presenting incorrect information about your company, products or claims.
3. Agent Transaction and Discovery
The third layer is where AI search starts to become much more interactive.
Rather than simply reading information and reporting it back to a user, an AI agent may eventually need to perform tasks on the user’s behalf.
That could include making a booking, purchasing a product, accessing account information or completing another transaction.
For businesses, this means their websites may need to expose certain capabilities in ways that AI agents can discover and securely use.
This is a relatively new area, but it could become increasingly important as consumers start using AI agents as an alternative way of interacting with businesses.
How the 50 Websites Performed
The audit examined 50 large websites across sectors including retail, SaaS, travel, publishing and financial services.
The websites were tested using an instrumented browser capable of capturing live HTTP responses, rendered DOM content, raw server HTML and machine-discovery endpoints. The testing was carried out on June 12, 2026.
The research focused on 12 established technical signals. Each was scored from zero to two, with the final result converted into a percentage.
Airbnb recorded the highest overall score at 79.2%.
Across the entire group, the average score was 56.6%, while the median was 58.3%. Only 10 of the 50 websites scored below 50%.
At first glance, an average above 50% may seem encouraging. However, the results look very different once the three layers are examined separately.
Retrievability Performs Best
The first layer achieved an average score of 74.4%, making it by far the most mature area of AI readiness.
Eight established elements were assessed, including:
- Robots.txt and AI user-agent directives
- Accessibility tree integrity
- ARIA labels and descriptive names
- Semantic HTML and document structure
- Efficient DOM density
- Server-rendered and clean HTML
- Machine-friendly forms and inputs
- Sitemap declarations
This layer has significant overlap with established technical SEO practices, which helps explain why websites performed comparatively well.
Most businesses have already spent years making their websites accessible to search engines, so some of the technical foundations required for AI systems are already in place.
Only three of the 50 websites scored below 50% in this area.
However, simply making content easy to retrieve does not guarantee that an AI system will interpret it correctly.
Attribution and Meaning Remains a Weak Point
Performance dropped considerably at the second layer, which recorded an average score of just 38.5%.
The established signals examined here were:
- JSON-LD schema and semantic richness
- Content Signals Policy
One positive finding was the widespread use of structured data. JSON-LD was detected on the homepage of 35 out of 50 websites, or 70% of the sample.
Almost all of those sites received the maximum score for this element.
However, that still means almost one-third of the websites had no JSON-LD structured data detected on their homepage.
Schema can give AI systems valuable context about what they are looking at. It can help establish that something is a product, identify its price and connect it to the correct brand, for example.
Without those signals, AI may still attempt to understand the page, but it has to rely more heavily on interpretation. That increases the possibility of incorrect conclusions and inaccurate AI-generated responses.
Content Signals Are Rare
The audit also found that only five of the 50 websites had implemented Cloudflare’s Content Signals Policy.
This provides instructions within robots.txt covering how crawlers can use website content for purposes such as search indexing, real-time AI responses and model training.
This gives businesses more control than simply choosing between allowing or blocking AI completely.
Instead, companies can make more specific decisions about how their content should be used.
The low adoption rate suggests that this is still an area where many SEO and marketing teams have considerable room to develop their approach.
Agent Readiness Is Still in Its Early Stages
The biggest gap appeared in the third layer.
Agent Transaction and Discovery recorded an average score of just 2.1%.
Only two of the 13 elements assessed could currently be classified as established:
- OAuth Discovery through an Authorisation Server
- OAuth Protected Resource Metadata
This low score is not particularly surprising. Agentic browsing and AI-powered commerce are relatively new developments, having emerged more prominently from late 2025 onwards.
Of the 48 websites where endpoint testing could be completed, 46 scored zero for the relevant tests.
Airbnb and Vercel had implemented OAuth authorisation server metadata, but neither had implemented protected resource metadata. As a result, both received only half of the available score for the relevant area.
These technologies are important because they can provide AI clients with information about how they can securely access a website’s services or APIs.
MCP and AI Commerce Could Change Things Quickly
Although the third layer is currently immature, it is also one of the areas most likely to develop rapidly.
Several emerging protocols could eventually make it much easier for AI agents to interact directly with business systems.
The Model Context Protocol, or MCP, is one example. It can allow AI systems to connect with servers and retrieve information directly.
Instead of an AI having to search through multiple product pages to work out availability or pricing, a direct connection could potentially provide information from a company’s live systems.
This could make the process faster while also reducing the risk of AI relying on outdated or incorrectly interpreted information.
Ecommerce is another area to watch.
Protocols such as Google’s Universal Commerce Protocol (UCP) and OpenAI’s Agentic Commerce Protocol (ACP) are designed around the idea of allowing AI systems to facilitate transactions within AI conversations.
If these technologies become widely adopted, websites that prepare early could gain an advantage over competitors that wait until the standards become mainstream.
A Low AI Readiness Score Does Not Always Mean Poor Preparation
There is an important caveat to the results.
A website receiving a low score does not automatically mean the company has neglected AI.
Different businesses have very different reasons for deciding how AI systems should interact with their content.
Some publishers, for example, have deliberately chosen to restrict AI crawlers.
The BBC, CNN and The Guardian are among the websites that block many AI bots. For publishers whose business models depend on attracting visitors directly to their websites, limiting AI access can be a perfectly rational commercial decision.
Amazon provides another interesting example.
It recorded an overall score of just 29.2% and was one of only three websites to score below 50% for retrievability.
That does not mean Amazon has failed to consider AI.
The company has deliberately introduced AI crawler directives designed to prevent most AI bots from accessing its website. If that is part of its wider strategy, implementing additional AI-readiness measures may not make commercial sense.
Other companies have taken a more selective approach.
Airbnb and Cloudflare have implemented access rules covering many major AI bots without completely blocking them. eBay and Tripadvisor have taken a more targeted approach by allowing some AI systems while restricting others.
The important point is that these companies have made deliberate decisions.
Most Websites Are Still Leaving AI Access to Chance
The more concerning finding is the number of businesses that appear not to have made a decision at all.
Of the 50 websites audited, 29 appeared to have no deliberate policy governing AI agent access.
Their websites did not clearly block AI systems, but they also did not explicitly allow or manage them through dedicated rules.
This effectively leaves the question of which AI systems can access their content to default settings and the behaviour of individual crawlers.
It also means businesses may have little control over how AI systems access, interpret and ultimately use their information.
For companies that want to benefit from AI search, that is an important gap to address.
AI Signals Are Not Guaranteed to Work
Another technology attracting attention is llms.txt.
The file is intended to provide AI systems with a simplified, human-readable guide to a website, including information about its structure and important content.
The audit classified llms.txt as a Frontier technology because there is currently no ratified industry standard or agreed governing body behind it.
Nevertheless, 11 of the 50 websites had published an llms.txt file.
That suggests there is growing interest in giving AI systems a clearer summary of a company’s website and content.
However, businesses should not assume that adding llms.txt or updating robots.txt will solve all their AI visibility problems.
AI crawler instructions are essentially requests rather than technical enforcement mechanisms. Major crawlers such as GPTBot, ClaudeBot, Google-Extended and Applebot-Extended generally respect these instructions, but compliance is not universal.
Content Signals face similar limitations.
Meanwhile, llms.txt currently represents more of an indication of intent than a guaranteed technical mechanism.
Expedia Shows Why One Signal Is Not Enough
Expedia provides a useful example.
The company has published an llms.txt file containing information designed to help AI understand its identity, content and capabilities.
However, Expedia’s overall score in the audit was just 33.3%, placing it among the lower-scoring websites in the group.
The website had no detected JSON-LD structured data, no sitemap declaration and only around a quarter of its content was delivered server-side.
The lesson is not that llms.txt is useless.
Instead, it shows why businesses should avoid treating any single AI optimisation technique as a solution in itself.
Just as traditional SEO relies on multiple signals, AI optimisation is likely to require a combination of technical accessibility, structured information, clear permissions and machine-friendly functionality.
SEO Teams Need to Think Beyond Citations
Being mentioned or cited by an AI system is only the starting point.
A brand can appear regularly in AI-generated answers and still have problems if the information being presented is inaccurate, incomplete or misunderstood.
The next stage of AI search is therefore less about simply getting visibility and more about controlling how AI accesses, understands and uses your information.
Different businesses may reach different conclusions about what that should look like.
Some may want maximum AI access. Others may prefer to restrict crawlers. Some may want specific AI systems to access their content while blocking others.
There is no single strategy that will suit every website.
What matters is making that decision deliberately.
The Next Stage of AI Search
The audit highlights a clear divide between traditional technical SEO and the requirements emerging around AI.
Most major websites have made reasonable progress with retrievability. Far fewer have invested in helping AI establish meaning, ownership and context. Almost none have yet built the infrastructure required for sophisticated AI agents to interact with their websites.
That is likely to change as agentic search and AI commerce develop.
For SEO teams, the opportunity is to start looking beyond rankings, traffic and citations. Technical decisions involving robots.txt, structured data, accessibility, APIs, authentication and machine-readable content could increasingly influence how a brand performs in AI-driven search.
AI is already accessing vast amounts of website content.
The question for businesses is no longer simply whether AI can find their website. It is whether AI can understand it accurately, respect how the information should be used and eventually interact with the business itself.
Those are technical challenges, but they are also opportunities for companies that start preparing before these standards become commonplace.
More Digital Marketing BLOGS here:
Local SEO 2024 – How To Get More Local Business Calls
3 Strategies To Grow Your Business
Is Google Effective for Lead Generation?
How To Get More Customers On Facebook Without Spending Money
How Do I Get Clients Fast On Facebook?
How Do You Use Retargeting In Marketing?
How To Get Clients From Facebook Groups