Product methodology
How the Search and AI Visibility check works
DomainOptic performs a point-in-time review of public page and robots.txt signals. The optional Important Pages map is a beta that compares a homepage with up to two selected same-site pages. This page records what is checked, how results should be interpreted, and what the check cannot establish.
Last reviewed against the linked primary sources: July 26, 2026.
What we request
The checker requests the submitted public homepage and the robots.txt file for the final homepage host. Redirects are bounded and restricted to the same registrable domain.
When the Important Pages map is opened, DomainOptic checks at most five same-site sitemap files. That bound supports one sitemap index plus four section sitemaps while preventing an unusually deep sitemap tree from creating an open-ended crawl. Sitemap URLs are selection candidates only. The user chooses up to two additional pages, and the same passive page analysis runs against each selected path. The map does not alter the saved homepage monitoring baseline.
It does not bypass authentication, paywalls, robots.txt, or server restrictions. It does not impersonate a vendor crawler or verify crawler IP addresses.
Signals reported
Homepage indexing directives
We read the initial homepage response, HTML robots metadata, and X-Robots-Tag headers for general, Googlebot, and Bingbot noindex directives.
Canonical and descriptive metadata
We report the canonical URL, title, meta description, page language, and whether JSON-LD is present. Presence does not mean the structured data is valid or eligible for a search feature.
Initial HTML readability
We measure text present before client-side JavaScript runs. This is an observation about the fetched response, not a content-quality or ranking score.
Selected robots.txt policies
We evaluate the published policy for selected current search crawlers, ads-verification crawlers, user-requested retrieval agents, training crawlers, AI-use controls, and a public web-corpus crawler. Every reported token links to the first-party source and its verification date.
Agent conventions (emerging)
We report declared Content-Signal preferences and whether the homepage answers an Accept: text/markdown request with Markdown. These are adoption observations, not search or AI requirements.
How crawler policies are grouped
Search discovery
Googlebot, Bingbot, Applebot, OAI-SearchBot, Claude-SearchBot, PerplexityBot, Amzn-SearchBot, MistralAI-Index, and Kimi-SearchBot.
Ads landing-page review
OAI-AdsBot is reported separately because OpenAI documents it for ChatGPT Ads landing-page validation. Allowing it does not imply that OAI-SearchBot is allowed.
User-requested retrieval
Claude-User, MistralAI-User, and Amzn-User are reported because their vendors publish robots.txt controls for them. User-triggered agents with unresolved or explicitly different behavior are not assigned a misleading allow or block result.
Training and AI-use controls
GPTBot, ClaudeBot, Meta-ExternalAgent, Amazonbot, KimiBot, Google-Extended, and Applebot-Extended. These controls are not interchangeable with search inclusion.
Public web corpus
CCBot, the crawler associated with the Common Crawl corpus.
A robots.txt result describes the policy we fetched. It does not prove that a genuine crawler can reach the site through a CDN, firewall, or bot-management layer.
Model names are not crawler names. If a major AI provider has not published a first-party robots.txt token and purpose, it is not shown. Qwen, DeepSeek, and Grok are not currently assigned guessed allow or block badges.
Review the AI Crawler Registry for every reported token, purpose, first-party source, and verification date.
How account monitoring works
Monitoring repeats the bounded homepage and robots.txt review three times per week. The first successful run establishes a silent baseline. A failed or unreadable homepage request preserves the last successful baseline and does not create a visibility-change email.
Later successful runs compare crawler-policy access, robots.txt reachability, homepage noindex directives, sitemap declarations, canonical host, initial HTML word count, and structured-data types. A sharp readability reduction must persist across two scheduled checks before it is alertable.
How robots.txt suggestions stay conservative
A crawler-specific user-agent group does not inherit the wildcard group. A short suggestion with only a crawler name and Allow: / can therefore drop wildcard restrictions for private or operational paths.
DomainOptic provides copyable text only when the checked homepage is blocked by a simple wildcard Disallow: /. The suggestion carries every other wildcard Allow and Disallow rule into the crawler-specific group, validates that the homepage becomes accessible, and limits the number of copied rules.
Existing crawler-specific groups, non-homepage blocks, large policies, overlapping patterns, and groups with additional directives stay manual. DomainOptic does not modify the scanned site. The site owner must review the full file and preserve restrictions for account, admin, API, private, and staging paths.
What the result cannot prove
- That a page is indexed, ranks for a query, or will appear in an AI-generated answer.
- That structured data is valid, accurate, or eligible for a search presentation.
- That unselected pages share the configuration found on the homepage or selected map pages.
- That a crawler is allowed by server, CDN, firewall, or account-level controls.
- That allowing a training crawler is required for search discovery or citations.
Primary sources
Crawler names and interpretations are maintained from first-party documentation. Vendor behavior can change, so the review date matters.
Use the observations as a starting point
Correct clear technical gaps, then confirm search behavior in the relevant webmaster tools. Content quality, evidence, reputation, query relevance, and each platform's systems remain outside this checker.
Run the visibility check or read the separate security scanner methodology.