State of Indian SEO

We Audited 226 Websites for AI Search Readiness

Jinto JosePublished 4 Aug 2026 13 min read

We Audited 226 Websites for AI Search Readiness

What's YOUR site's SEO score?

Free scorecard in 30 seconds. No signup, no jargon.

Prefer email? Get free, plain-English SEO tips in your inbox:

Most websites aren't ready for AI search, and the reason is duller than the advice suggests. We ran our free audit across 226 real small-business websites between 4 June and 3 August 2026. The average score was 75 out of 100, which sounds fine. But heading structure came out as the single worst category at 54 out of 100, 58% of sites have headings that skip levels, and 16% have no main heading at all. Before anyone worries about llms.txt, most sites can't be read cleanly enough to be quoted.

That last point is the finding. AI engines answer questions by lifting passages out of pages they can parse. A page with no clear structure gives them nothing clean to lift.

If you want your own site's numbers before reading ours, run the free SEO audit. It's the same 30-second check these 226 sites went through, no signup.

What we measured, and on how many sites

We audited 226 unique domains. The raw number was 244 completed audits; 18 were repeat audits of a site we'd already seen, and counting a site twice would have skewed every percentage, so each domain votes once. The audits ran between 4 June and 3 August 2026.

These are ordinary business sites, not big brands. By domain ending: 152 .com, 24 .in, 14 .com.au, 11 .co.uk, with the rest spread across .net, .io, .ca and others. So it's a global sample with an Indian and Australian tilt, not an India-only one.

Two honest limits before any of the numbers, because they change how you should read them. First, these sites weren't picked at random. Someone chose to run each one through a free SEO audit, which selects for people who already suspected something was wrong. Treat this as a portrait of sites whose owners are paying attention, not of the whole web. Second, the audit reads a site's homepage, so everything below is a statement about homepages.

That second limit is why we cut two of our most eye-catching numbers from this post. Just over half the sites had no author or byline, and 44% had no published or updated date. Both matter enormously for whether AI engines trust a page. But homepages legitimately don't carry bylines or datestamps, so quoting those figures here would be measuring the wrong thing on the wrong page. They belong in a study of article pages, which isn't this one.

Where do sites actually score worst?

Structure, by a wide margin. Averaged across all 226 sites, the six categories came out like this:

CategoryAverage score
Security96.2
Mobile88.5
Content79.4
Performance78.0
Meta tags64.2
Headings54.0

Security at 96 and mobile at 88 tell a clear story: the industry spent a decade telling people to move to HTTPS and fix their phone layouts, and people did it. Those campaigns worked.

Headings at 54 is ten points below the next worst category. Nobody ran a campaign about heading structure, because for most of Google's history it barely mattered. A page could have three <h1> tags and a heading order that jumped from <h2> to <h5> and still rank perfectly well, because Google's ranking systems were tolerant of messy markup. AI answer engines are less forgiving, since the heading tree is one of the main ways they work out which chunk of a page answers which question.

The overall average score was 75.3 and the median was 76, on a range from 48 to 98. A quarter of the sites (57 of 226) had at least one issue we class as critical.

The eight findings, in order

Each percentage below is the share of sites where we found that problem. Where a check was added to the audit partway through the study, we've said so and measured it only against the sites audited after it existed, which is the honest denominator.

58% have headings that skip levels. The page jumps from a main heading straight to a sub-sub-heading, or wanders back and forth. To a person it looks fine, because visual size carries the meaning. To a parser the outline is broken.

51% have images with no description. Alt text is the only thing a machine can read on an image. Half these sites have images carrying real information that's invisible to anything that isn't a human eye.

50% have no llms.txt file. Measured on the 205 sites audited after we added this check on 21 June. This is the one everybody writes about, and it's genuinely worth adding, but note where it sits in this list. It's an emerging standard that not every engine reads yet, and a site with broken headings gains little from publishing one. If you want the detail, we wrote a plain-English guide to what llms.txt is and whether you need one.

42% load a lot of images on a single page, which is a speed problem before it's anything else.

36% have no link previews set up. No Open Graph tags, so when someone shares the site on WhatsApp or LinkedIn it renders as a bare grey link.

31% have no question-style headings. Measured on the 164 sites audited after we added this check on 29 June. If nobody on your site ever writes a heading in the form a customer would actually ask, there's nothing for an engine to match a question against.

22% have no structured data at all. No schema markup telling machines what kind of business this is.

21% have more than one main heading, and 16% have none. That second figure is the one that stops us short. Thirty-seven sites out of 226 have no <h1> element anywhere on the homepage, which means there is no machine-readable statement of what the page is about.

You can check the last two on your own site in about a minute with our free schema markup generator, and check whether AI crawlers can reach you at all with the AI crawler checker.

How this compares to the other studies

Two other groups are running primary research on AI readiness right now, and both are worth reading. They also both measure something different from what we measured, which is why we bothered publishing this.

The AI Visibility Index from Website Auditor covers 531 audits across 458 domains and 38 sectors, running from 30 March to 2 August 2026. Its headline is stark: 94.8% of the sites it studied never appear when an AI assistant answers a buying-intent question about their category, and only 1.9% of answers named the business being asked about. On the technical side it reports that 79.4% publish a robots.txt, 54.6% publish schema.org data, 19.3% publish LocalBusiness schema specifically, and 8.9% block at least one AI crawler.

AI Visibility's research programme crawls roughly 1,995 domains each quarter, drawn from the Tranco list's global and UK top 1,000. It measures adoption of AI discovery files like llms.txt, ai.txt and identity.json, whether those files actually validate, and how sites use robots.txt to manage AI crawler access.

Neither study measures heading structure, <h1> presence, alt text, bylines or dates. The first measures the outcome, which is whether you get named. The second measures AI-specific files on the web's largest domains, which are big brands with dedicated teams. The gap between them is the ordinary page structure that decides whether a passage can be extracted in the first place, on the kind of small business site that has no dedicated team at all. That's the gap this data fills, and it suggests the 94.8% figure has a duller cause than most of the advice assumes.

Read alongside each other, the three datasets point the same way. Website Auditor found 54.6% publishing schema and we found 78% publishing schema, and the difference is almost certainly sample composition rather than a contradiction, so don't read too much into either number alone.

What should you actually fix first?

Fix the heading tree, then the alt text, then everything else. This is the opposite of the order most AI-visibility advice gives, and the data is why.

Start by making sure every page has exactly one <h1> that states what the page is about, and that the headings below it descend in order without skipping. In WordPress, Wix, Shopify and Webflow this is a matter of picking the right heading style from a dropdown rather than picking the size you liked the look of. It's genuinely an afternoon of work on a small site, it costs nothing, and it helps ordinary Google rankings and screen-reader users at the same time.

Then write alt text on any image that carries information. Then, once the structure is sound, the AI-specific work is worth doing: add schema, publish an llms.txt, make sure you aren't blocking GPTBot or Google-Extended. We've written the full ordered list in how to appear in Google AI Overviews, and the background on why any of this works in what generative engine optimization actually means.

The honest caveat on all of it: fixing your headings won't get you cited by ChatGPT tomorrow. Being named in an AI answer still depends heavily on whether other sites talk about you, which is slow and mostly outside your control. What structure does is remove the mechanical reason you'd be skipped. It's necessary and it isn't sufficient.

Methodology, in enough detail to check us

If you're going to cite this, you should be able to see exactly how each number was produced and where it stops being safe to generalise.

Instrument. Every figure comes from the public RankAgent audit — the same one linked above, running the same checks in the same order on every site. Nothing was hand-coded, hand-scored or re-run to get a better number. You can put any site through the identical instrument yourself, which is the practical version of reproducibility here.

Window and unit. Audits completed between 4 June and 3 August 2026. 244 completed audits, deduplicated to 226 unique registrable domains (where a domain was audited more than once we kept the most recent run). The unit of analysis is one homepage per domain.

Sampling. Self-selected, not random. Each site is one somebody chose to run through a free SEO audit. That biases toward owners who already suspected a problem, and away from sites nobody worries about — in both directions. There is no control group, no weighting and no attempt to match the sample to any population, so treat every figure as descriptive of this sample rather than as an estimate for "the web".

Scope. Homepages only. Any finding here is a statement about homepages and cannot be extended to article or product pages without measuring those separately — which is why the byline and datestamp figures were cut from this post rather than reported.

Per-check denominators. Checks added partway through the window were measured only against sites audited afterwards, so the denominator varies by check and is stated inline every time. Nothing is reported against a denominator it wasn't measured on.

StatisticDenominatorNote
Average score 75 / median 76 / range 48–98226Whole sample
Headings category 54.0226Whole sample
58% headings skip levels226Whole sample
21% multiple H1 · 16% no H1226Whole sample
78% publishing schema226Whole sample
50% no llms.txt205Check added 21 June 2026
31% no question-style headings164Check added 29 June 2026

Rounding. Percentages are rounded to the nearest whole number, so a column of them will not always total exactly 100.

What we measured, and what we did not. We measured markup: what is present in a homepage's served HTML, and how our audit scores it. We did not measure AI citations, mentions, referral traffic or rankings for any of these 226 sites. So the central claim of this post is a claim about mechanics — a page with a broken heading tree is harder to extract a passage from — and not a demonstration that a specific site was omitted from an AI answer because of its headings. Nothing here establishes causation, and a missing llms.txt or a skipped heading level is not evidence of why any engine left a site out. The honest form of the argument is the one in the section above: structure is necessary, and it is not sufficient.

Conflict of interest. RankAgent sells software that fixes the problems this data describes. The instrument is our own product, and its category weights are our methodology rather than measured ranking factors. Weigh the numbers accordingly, and reproduce anything that matters to you.

Frequently asked questions

Are websites ready for AI search?

Mostly not, on the evidence here. Across 226 small-business sites we audited, heading structure scored 54 out of 100, 58% had headings that skip levels and 16% had no main heading at all. The AI-specific work most advice leads with matters less than that, because a page that can't be parsed cleanly can't be quoted regardless of what files it publishes.

What's a good SEO score?

In this sample the average was 75 and the median 76, across a range of 48 to 98, so a score in the mid-70s is genuinely ordinary rather than good. A quarter of sites carried at least one critical issue. Bear in mind these owners had chosen to run an audit, so the sample skews toward people already looking for problems.

Does llms.txt actually matter for AI search readiness?

It's worth adding and it isn't the first thing to fix. Half the sites we measured don't have one, but it's an emerging standard that not every AI engine reads today, and publishing one on a site with a broken heading tree doesn't fix the underlying problem. Sort out the structure, then add the file.

Why does heading structure matter more for AI than for Google?

Google's ranking systems have tolerated messy markup for years, because they lean on many other signals. AI answer engines pull specific passages out of pages to build an answer, and the heading tree is one of the main ways they identify which passage answers which question. A broken outline gives them nothing reliable to lift.

How was this study run?

We ran our standard free audit on each site's homepage, covering the same checks anyone gets from the public scorecard. We deduplicated by domain, so 244 completed audits became 226 unique sites, and where a check was added partway through we measured it only against sites audited after that date.

Can I check my own site?

Yes, free and without signing up. Run your site through the free SEO audit to get the same score and the same checks these 226 sites got, in about 30 seconds, in plain English rather than jargon.

What's YOUR site's SEO score?

Free scorecard in 30 seconds. No signup, no jargon.

Prefer email? Get free, plain-English SEO tips in your inbox:

JJ

Jinto Jose — Founder, RankAgent

Founder of RankAgent, AI SEO software for agencies and freelancers. I build the audit engine, run it on real client sites, and publish what it finds.

Keep reading

View all articles →