We Checked 80 Popular SaaS Tools for AI Crawler Access. 70 Block Nothing.
Published August 17, 2026 · By DomainOptic · 7 min read
The Question We Measured
The loudest story about AI crawlers in 2026 is refusal: publishers locking out GPTBot, lawsuits over training data, robots.txt as a picket line. We wanted to know whether the tools developers actually use every day behave that way.
So we measured it. On August 17, 2026 we fetched robots.txt from 80 well-known developer and SaaS products and parsed each file with the same parser that powers our Search and AI Visibility check and AI crawler registry. No sampling, no estimates: one dated snapshot of what each company publishes.
The Headline Numbers
- 80 domains checked; 77 had a readable robots.txt. Three (npmjs.com, miro.com, 1password.com) did not return a readable file to a generic fetcher, most likely bot management in front of the file itself.
- 70 of the 77 readable files (91 percent) block no documented AI crawler at all. Not GPTBot, not ClaudeBot, not PerplexityBot, not the training opt-outs.
- GPTBot is the most-blocked AI crawler, and it is blocked by 5 sites out of 77. That is 6 percent: Figma, Canva, Loom, Reddit, and Medium.
- ClaudeBot and CCBot are each blocked by 4 sites. Google-Extended by 2 (Figma and Reddit).
- Exactly one company blocks the answer-engine crawlers: Reddit blocks PerplexityBot, OAI-SearchBot, Claude-SearchBot, and essentially every other documented AI token, consistent with its licensing strategy.
The Pattern Inside the Blockers
The few sites that do block are selective in a telling way. Figma blocks GPTBot and Google-Extended, which govern model training, while leaving the search and on-demand retrieval crawlers alone. The same shape appears at Canva and Loom: training no, answers yes.
That split is the strategy the raw numbers suggest across the whole sample. Being quotable by ChatGPT, Copilot, and Perplexity is distribution. Donating training data is not. Most tool companies apparently want the first, and only a design-heavy minority spends rules on the second.
Reddit is the exception that proves it: when your content is the product being licensed, you block everyone and sell access at the door.
Why This Matters If You Run a SaaS Site
If you have been debating whether to block AI crawlers defensively because everyone seems to, the measured answer is that among the tools your own team uses, almost nobody does. The competitive default in this market is open. Blocking the retrieval crawlers means assistants cannot quote your docs or recommend your product when someone asks, and this sample suggests your competitors are not making that trade.
The reverse check matters just as much: hosting and CDN defaults, WAF rules, and bot management can block crawlers your robots.txt allows. Platforms like Vercel handle a lot for you but leave application-layer choices to you, which is why we keep a separate Vercel security checklist. To see what your own site currently tells every documented crawler, run the visibility check: it reads your robots.txt against the same registry used in this study.
Method, So You Can Reproduce It
- Date: August 17, 2026, single snapshot.
- Sample: the 80 domains listed below, chosen as widely used developer and SaaS tools before any data was collected. No domain was added or removed after measurement.
- Fetch: https robots.txt for each domain (www fallback), generic research user agent, redirects followed.
- Parsing: the same robots.txt parser behind our public checker, evaluated against every crawler token in our documented registry, which records only tokens with published vendor documentation.
- Caveats: robots.txt is published policy, not enforcement. A site can block at the network layer while allowing in robots.txt, and vice versa. Three unreadable files are reported as unreadable, not assumed. This is one day; policies change.
The sample: vercel.com, netlify.com, github.com, gitlab.com, stripe.com, notion.so, figma.com, linear.app, slack.com, zoom.us, dropbox.com, airtable.com, asana.com, trello.com, monday.com, clickup.com, intercom.com, zendesk.com, hubspot.com, salesforce.com, mailchimp.com, sendgrid.com, twilio.com, cloudflare.com, digitalocean.com, heroku.com, render.com, railway.app, supabase.com, mongodb.com, redis.io, postman.com, docker.com, kubernetes.io, circleci.com, npmjs.com, pypi.org, rubygems.org, wordpress.org, shopify.com, wix.com, squarespace.com, webflow.com, framer.com, canva.com, miro.com, loom.com, calendly.com, typeform.com, surveymonkey.com, grammarly.com, 1password.com, bitwarden.com, okta.com, auth0.com, twitch.tv, discord.com, reddit.com, medium.com, substack.com, ghost.org, buffer.com, hootsuite.com, zapier.com, ifttt.com, make.com, retool.com, bubble.io, plausible.io, posthog.com, basecamp.com, gumroad.com, lemonsqueezy.com, paddle.com, chargebee.com, algolia.com, meilisearch.com, sentry.io, datadoghq.com, newrelic.com.
Questions about the data or want your category measured next? The parser and registry are public parts of the product, so you can check any single site yourself in seconds.
Check what your site tells AI crawlers
Search Visibility
Continue with visibility evidence
Scan your domain