Agent Web Index › hexun.com

Can AI assistants read hexun.com?

Measured on 2026-09-19 by asking the site 7 times — once as a browser, once as each of the 6 crawlers that feed ChatGPT, Claude, Perplexity, Gemini, Meta AI, Apple Intelligence and Doubao — and comparing what came back. The AI-training opt-out tokens that never crawl are read from robots.txt instead. Tranco rank #5,592.

D50 / 100
2 of 6 crawlers can read this page.
More readable than 7% of the 47,934 sites measured so far. No JSON-LD.

What each crawler got back

CrawlerHTTPResultrobots.txt
ClaudeBot (Claude)200 can read it no-file
GPTBot (ChatGPT)200 can read it no-file
OAI-SearchBot (ChatGPT Search)567 error no-file
PerplexityBot (Perplexity)567 error no-file
Google-Extended (Gemini, AI Overviews)
a robots.txt token, not a crawler: it controls how already-crawled pages may be used, and never makes a request of its own
not opted out no-file
Meta-ExternalAgent (Meta AI) no-answer no-file
Amazonbot (Alexa, Rufus)567 error no-file

What to change, in order

  1. Let in the 4 crawlers your own robots.txt already allows+4 crawlers
    OAI-SearchBot, PerplexityBot, Meta-ExternalAgent, Amazonbot are turned away before reading the page (answered with an HTTP error; the connection never completed), while robots.txt permits them — so this block is not written in your site. No CDN signature was found in the response headers, so the refusal comes from the origin server itself or from a WAF this index does not recognise. Fixing it takes this domain from 2 to 6 of 6 crawlers.
  2. Declare the facts in JSON-LD
    There is no JSON-LD on the page, so every fact — who you are, what you sell, the price — has to be guessed out of the prose and the layout. A schema.org block is the difference between being quoted correctly and being paraphrased.
  3. Publish sitemap.xml and llms.txt
    sitemap.xml and llms.txt are missing, so an assistant has to discover the site by following links. llms.txt is the emerging convention for telling an assistant which pages actually matter.
  4. Fix the plain structure: one <h1>, a title, a description, alt text
    The structure check scores 28/100. These are the cheapest signals on the page and the first ones an assistant uses to decide what the site is.

Get told if this changes

One email only when a measured crawler flips on hexun.com, served to refused or back. No schedule, no newsletter; double opt-in, one-click stop.

The checks

History

2026-09-13: 27 · 2026-09-14: 42 · 2026-09-15: 57 · 2026-09-16: 47 · 2026-09-17: 42 · 2026-09-18: 52 · 2026-09-19: 50

Full AI-visibility report for hexun.com → Scan your own site →

Sites with a similar score

ledevoir.comD 55weserv.nlD 51hec.caD 50hillsdale.eduD 50hankooki.comD 50idealmedia.ioD 50footwearnews.comD 50kering.comD 50pornorama.comD 50webdomaine.caD 49oercommons.orgD 48ninjakiwi.comF 42

Browse the whole index →

Embed this score

Put the badge on hexun.com — it links back here, and re-measures every time this index re-crawls.

hexun.com AI readability: D 50/100
<a href="https://aivis.lumnika.com/ai-readiness/hexun.com"><img src="https://aivis.lumnika.com/ai-readiness/hexun.com/badge.svg" alt="AI readability"></a>
Method. 7 live HTTP requests (one per crawler, one as a browser) plus robots.txt, llms.txt, sitemap.xml and security.txt, 12-second timeout each, from a single vantage point. Blocking AI crawlers is a legitimate choice, not a failure: this page records what is true, not what should be. Domain comes from the Tranco research list. One page per domain — the homepage — is audited.

Part of the Agent Web Index, 47,934 domains measured, updated continuously.