197 URLs Discovered, 100 Indexed: A July 2026 Search Console Case Study
Published July 23, 2026 · By DomainOptic · 6 min read
What the Numbers Actually Mean
On July 19, 2026, Google Search Console successfully read DomainOptic's sitemap and reported 197 discovered pages. The Page Indexing report, last updated July 9, showed 100 indexed pages and 156 not indexed pages across all known URLs.
Those numbers do not form a clean 100-out-of-197 indexing rate. They come from different reports, scopes, and snapshot dates. The 197 belongs to the current submitted sitemap. The 256 total in the Page Indexing report includes older redirects, retired routes, parameter variants, and other URLs Google knows about.
That distinction matters because "discovered" means Google found a URL through a sitemap. It does not mean Google selected the page for indexing.
The First-Party Snapshot
The authenticated DomainOptic Search Console property showed:
- Sitemap status: Success, read July 19, 2026.
- URLs discovered from the submitted sitemap: 197, read July 19, 2026.
- Indexed pages across all known URLs: 100, report updated July 9, 2026.
- Not indexed pages across all known URLs: 156, report updated July 9, 2026.
- Crawled, currently not indexed: 120, report updated July 9, 2026.
- Page with redirect: 21, report updated July 9, 2026.
- Excluded by noindex: 11, report updated July 9, 2026.
- Alternate page with proper canonical: 4, report updated July 9, 2026.
The most important number was not 197. It was the 120 URLs Google had already crawled but had not selected for indexing.
What Was in the Crawled but Not Indexed Group
We classified the 120 example URLs from Search Console:
- Glossary terms: 65 URLs.
- Blog articles: 35 URLs.
- Technology security guides: 7 URLs.
- Core pages, publisher files, retired routes, and URL variants: 13 URLs.
This was not primarily a crawl-access failure. Google had fetched the pages.
A live URL inspection on one glossary page made the point clearer. Google reported that the page could be indexed, crawl was allowed, the fetch succeeded, and the user-declared canonical matched Google's selected canonical. The page still was not indexed in the stored result.
Why Another Sitemap Submission Was Not the Main Fix
Google's July 2026 sitemap documentation describes sitemap submission as a hint and says it does not guarantee crawling or indexing. It also says a sitemap should contain the canonical URLs a site wants in search results.
The sitemap was doing its discovery job. Repeatedly submitting the same URLs would not give the unselected pages more distinct value.
Google's July 2026 AI optimization guide makes the same point from another direction: a high quantity of pages does not make a site higher quality or more relevant. The guide emphasizes useful, original, non-commodity content and warns against creating separate pages for query variations.
The Internal-Link Finding
We also measured the crawler-visible static link graph generated by the site. This was a local HTML analysis, not Google's rendered link graph.
- 42 of 44 blog article URLs had two or fewer incoming static references.
- 15 of 16 technology security URLs had two or fewer incoming static references.
- 26 glossary URLs had two or fewer incoming static references.
This did not prove why Google excluded a page. It did show that many pages were connected mainly through a hub rather than through useful, contextual pathways from related content.
What We Changed
We chose measurement and content clarity over URL growth.
- Split the sitemap by purpose. The root sitemap is now an index for core pages, blog articles, glossary terms, and security guides. Google does not need this split for a 198-page site, but its current documentation says separate sitemaps can be useful for tracking section performance in Search Console.
- Kept every established route stable. Canonicals and trailing-slash URLs did not change.
- Strengthened contextual pathways. Hubs now point people toward a smaller set of starting resources, and related articles connect to the relevant tool or methodology instead of relying on a large global link list.
- Removed obsolete FAQ structured data. Visible questions remain useful to readers, but Google retired FAQ rich results in May 2026, so the site no longer publishes the markup as if it created a current Google search feature.
- Published this case study. It is first-party evidence with dates, scope, and limitations rather than another generic indexing checklist.
- Added validation gates. The build now rejects duplicate sitemap URLs, missing child sitemaps, stale section dates, canonical mismatches, noindex URLs in sitemaps, and malformed structured data.
What We Did Not Do
- We did not add pages merely to increase the discovered count.
- We did not mass-request indexing for 120 URLs.
- We did not add special AEO or GEO markup.
- We did not claim that an allowed crawler guarantees an AI citation.
- We did not delete established glossary routes from one stale report.
Google says its generative search features still depend on normal crawlability and indexing. It also says that \llms.txt\ and special AI markup are not required for Google Search. Bing's February 2026 AI Performance announcement recommends measuring actual citations, cited pages, and grounding queries rather than treating crawler access as proof of inclusion.
How We Will Judge the Result
The first valid comparison starts only after Search Console's Page Indexing report advances beyond its July 9 snapshot and reads the new section sitemaps.
We will look for:
- the indexed share of each sitemap section;
- whether improved pages leave the crawled-but-not-indexed group;
- impressions and clicks for the routes that already demonstrate demand;
- actual citation activity in Bing AI Performance;
- Google generative-search visibility if the new report becomes available for the property.
There is no honest guarantee that the indexed count will rise. The useful outcome is better evidence about which content families Google selects, plus fewer weak signals competing with the site's strongest tools and original material.
Reproducibility and Limits
The public study data file records the report dates, counts, page-family classification, local static-link measurements, and source URLs used here. Search Console observations require property access and cannot be independently queried by a visitor. The local link analysis can be repeated from the public HTML, but it does not represent Google's private rendering or ranking systems.
Use Search and AI Visibility to inspect public technical signals on your own site, then confirm real indexing and citations in the webmaster tools that own those measurements.
Search Visibility
Continue with visibility evidence
Scan your domain