Product methodology
How the Search and AI Visibility check works
DomainOptic performs a point-in-time review of public homepage and robots.txt signals. This page records what is checked, how results should be interpreted, and what the check cannot establish.
Last reviewed against the linked primary sources: July 20, 2026.
What we request
The checker requests the submitted public homepage and the robots.txt file for the final homepage host. Redirects are bounded and restricted to the same registrable domain.
It does not bypass authentication, paywalls, robots.txt, or server restrictions. It does not impersonate a vendor crawler or verify crawler IP addresses.
Signals reported
Homepage indexing directives
We read the initial homepage response, HTML robots metadata, and X-Robots-Tag headers for general, Googlebot, and Bingbot noindex directives.
Canonical and descriptive metadata
We report the canonical URL, title, meta description, page language, and whether JSON-LD is present. Presence does not mean the structured data is valid or eligible for a search feature.
Initial HTML readability
We measure text present before client-side JavaScript runs. This is an observation about the fetched response, not a content-quality or ranking score.
Selected robots.txt policies
We evaluate the published policy for selected current search crawlers, user-requested retrieval agents, training crawlers, AI-use controls, and a public web-corpus crawler.
How crawler policies are grouped
Search discovery
Googlebot, Bingbot, Applebot, OAI-SearchBot, Claude-SearchBot, PerplexityBot, Amzn-SearchBot, MistralAI-Index, and Kimi-SearchBot.
User-requested retrieval
Claude-User and MistralAI-User are reported because their vendors publish robots.txt controls for them. User-triggered agents with unresolved or explicitly different behavior are not assigned a misleading allow or block result.
Training and AI-use controls
GPTBot, ClaudeBot, Meta-ExternalAgent, Amazonbot, KimiBot, Google-Extended, and Applebot-Extended. These controls are not interchangeable with search inclusion.
Public web corpus
CCBot, the crawler associated with the Common Crawl corpus.
A robots.txt result describes the policy we fetched. It does not prove that a genuine crawler can reach the site through a CDN, firewall, or bot-management layer.
Model names are not crawler names. If a major AI provider has not published a first-party robots.txt token and purpose, it is not shown. Qwen, DeepSeek, and Grok are not currently assigned guessed allow or block badges.
What the result cannot prove
- That a page is indexed, ranks for a query, or will appear in an AI-generated answer.
- That structured data is valid, accurate, or eligible for a search presentation.
- That every page on the site shares the homepage configuration.
- That a crawler is allowed by server, CDN, firewall, or account-level controls.
- That allowing a training crawler is required for search discovery or citations.
Primary sources
Crawler names and interpretations are maintained from first-party documentation. Vendor behavior can change, so the review date matters.
Use the observations as a starting point
Correct clear technical gaps, then confirm search behavior in the relevant webmaster tools. Content quality, evidence, reputation, query relevance, and each platform's systems remain outside this checker.
Run the visibility check or read the separate security scanner methodology.