What to check

Meta robots, robots.txt, sitemap, canonical, language versions, server HTML and CDN responses to crawler user agents.

Different crawlers

Search crawlers and training crawlers can have different purposes. Access should be intentional, not copied from another robots.txt.

Why access is not enough

Technical access only allows the page to be read. Recommendation still needs a clear product and independent evidence.

Practical takeaway

Do not optimise for an imagined algorithm. Capture a baseline, identify a specific gap, then change sources, pages and mentions.