Is the FTSE 350 ready for AI search? We checked 322 corporate websites
On 23 September 2026 we read the robots.txt file, llms.txt file and homepage code of every FTSE 350 corporate website. The doors are open to AI. The introductions are missing.
Are FTSE 350 companies ready for AI search?
Mostly open, rarely introduced. Only 4 of 271 readable robots.txt files (1.5%) block an AI search crawler, so almost every FTSE 350 corporate site can be read and cited by ChatGPT, Claude and Perplexity. But only 103 of 295 readable homepages (35%) publish Organization structured data that tells an AI engine who the company is, and only 26 of 322 sites (8%) have an llms.txt file.
The doors are open. The introductions are missing.
322 corporate websites, scanned 23 September 2026. Almost all of them let AI search engines in. Most never tell those engines who the company is.
- block an AI search crawler
- 1.5% block an AI search crawler 4 of 271 readable robots.txt files
- never mention an AI crawler
- 92% never mention an AI crawler 250 of 271 robots.txt files are silent on AI
- say who the company is
- 35% say who the company is 103 of 295 homepages publish Organization data
- have an llms.txt file
- 8% have an llms.txt file 26 of 322 corporate websites
Does the homepage tell AI engines who the company is?
One square per corporate website. Hover a square to see the company.
- Publishes Organization data (103)
- No Organization data (192)
- Homepage could not be read (27)
Which AI crawlers do FTSE 350 sites block?
Number of corporate robots.txt files blocking each crawler, out of 271 readable. Bars run from 0 to 10 sites.
Is the FTSE 100 any further ahead?
Share of sites doing each thing, FTSE 100 against FTSE 250.
- FTSE 100
- FTSE 250
Publish Organization data
List official profiles (sameAs)
Name any AI crawler in robots.txt
Have an llms.txt file
What did we check?
We took the 350 companies in the FTSE 100 and FTSE 250 after the quarterly review of 18 September 2026 and found each one's main corporate website. Some companies share a site, mostly investment trusts run by the same fund manager, so the sample is 322 unique websites.
On 23 September 2026 we made one request to each of three public files per site: robots.txt (the file that tells crawlers what they may read), /llms.txt (a proposed file that summarises a site for AI tools) and the homepage itself. We read the homepage as a crawler does, without running JavaScript, and looked for Organization structured data: the block of code that states a company's name, legal name, logo and official profiles in a form machines can read.
271 robots.txt files and 295 homepages were readable. The rest refused automated requests or returned errors, and we report percentages against the readable total.
Do FTSE 350 companies block AI search engines?
Almost never. Only four of 271 readable robots.txt files (1.5%) block any of the crawlers that fetch pages for AI answers: OAI-SearchBot and ChatGPT-User for ChatGPT, Claude-SearchBot and Claude-User for Claude, and PerplexityBot and Perplexity-User for Perplexity. None of the FTSE 100 does.
The exceptions are specific. ITV plc's corporate site blocks ChatGPT's search crawler and OpenAI's training crawler. Morgan Sindall Group blocks PerplexityBot, with a comment in its robots.txt saying there is no benefit to allowing it. Ruffer Investment Company blocks ChatGPT's live fetcher and PerplexityBot. Paragon Banking Group goes furthest: its robots.txt blocks all 14 AI crawlers we checked, search and training alike.
For a listed company this matters. When an analyst, journalist or job candidate asks an AI assistant about the business, a blocked site cannot be quoted, so the answer is built from whatever else the engine can find.
Are they opting out of AI training?
Rarely, and mostly by default. Nine sites (3.3%) block at least one AI training crawler such as GPTBot, ClaudeBot, Google-Extended or CCBot. In the FTSE 100 that is JD Sports and M&G.
The bigger finding is silence. Only 21 of 271 robots.txt files (7.7%) mention any AI crawler by name. For the other 92%, the decision about AI training and AI search has been made by default rather than by anyone in the business. One company, Taylor Wimpey, publishes Cloudflare's newer Content Signals line, which states separately whether content may be used for search, AI answers and training.
Do they tell AI engines who they are?
Mostly not. Only 103 of 295 readable homepages (35%) publish Organization structured data or a close variant such as Corporation or NewsMediaOrganization. In the FTSE 100 it is 32 of 91 (35%).
Several of the largest companies publish some structured data but never describe the company itself. HSBC's homepage describes its website and a video. Barclays describes the web page. AstraZeneca describes its website and site search. London Stock Exchange Group publishes a breadcrumb trail. None of the four declares the organisation behind the site.
Fewer still connect the dots. Only 51 homepages (17%) list their official profiles, such as LinkedIn, Wikipedia or the company's London Stock Exchange page, in a sameAs field. That field is how a machine confirms that the HSBC on one site is the HSBC on another.
Why it matters: AI assistants build answers about a company from many sources, and they have to decide which facts belong to which entity. A clear, consistent Organization record on the official site is one of the few signals a company fully controls. Structured data added by JavaScript after the page loads is often invisible to AI crawlers, which is why we read the raw page.
How many FTSE 350 companies have an llms.txt file?
26 of 322 sites (8%). In the FTSE 100 they are Beazley, Croda International, Reckitt, Weir and Whitbread. Whitbread's file says it was generated automatically by an SEO plugin rather than written by hand.
That is well below the 28% adoption Ahrefs found across its own, more technical sample of 137,210 sites. We do not think the FTSE is missing much. As we set out in why llms.txt does almost nothing for AI search, the major AI engines rarely request the file. Organization data and an open robots.txt matter far more.
What should a listed company do about AI search?
Three things, in order. None needs new software.
1. Decide on AI crawlers on purpose. Keep the search crawlers (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, PerplexityBot) allowed on the corporate and investor site, then make a separate, recorded decision about training crawlers.
2. Publish one Organization record on the homepage, in the page's own HTML. Include the legal name, the trading name, logo, ticker symbol, registered address and a sameAs list pointing to the official LinkedIn page, Wikipedia and Wikidata entries, Companies House record and London Stock Exchange page.
3. Check the investor site too. Several groups keep their corporate site on a subdomain or a separate domain with its own robots.txt, and that is the site AI engines tend to quote on results, leadership and strategy.
What are the limits of this study?
It covers the main corporate domain only, not consumer brand sites, which often have different rules. Investment trusts managed by the same firm share one site. 51 robots.txt files and 27 homepages could not be read automatically. We read raw HTML, so structured data added only by JavaScript is not counted. robots.txt is a request, not a lock, and firewall rules that block AI crawlers outright are not measured here. Constituents change quarterly.
Can I use this data?
Yes. The full dataset, one row per website with every check we ran, is free to use under a Creative Commons Attribution licence (CC BY 4.0). Please credit We Are All Connected and link to this page. Download the CSV. We will rerun the scan every month and update this page.
What is an AI search crawler?
A bot that fetches web pages so an AI assistant can quote or cite them in an answer, for example OAI-SearchBot for ChatGPT search or PerplexityBot. It is different from a training crawler, such as GPTBot, which collects pages to train future models.
What is Organization structured data?
A short block of code, usually JSON-LD in the page's HTML, that states facts about a company in a standard form: name, legal name, logo, address and official profiles. Search engines and AI systems use it to recognise the company as one entity.
Does blocking GPTBot stop a site appearing in ChatGPT?
No. GPTBot is OpenAI's training crawler. ChatGPT search uses OAI-SearchBot and ChatGPT-User. A site can block training and still be cited, which is the setup five FTSE 350 sites use.
Want to make this practical?
Book a call and we will walk through your current stack, what is worth wiring first, and what should wait.
Book a call