Skip to main content

Insight · Crawling, Indexing & Canonicals

Found, crawled, indexed: Clearly distinguishing the three states

Discovery, retrieval, and indexing are separate steps with different causes and verification methods. A URL can remain in place at any stage.

For SEO professionals and web developers, "Differentiating Between Found, Crawled, and Indexed" highlights the difference between "source of evidence" and "time reference." "Crawl equals index" is a typical warning sign.

Published: 3 min read · Author:

What is the difference between the states of found, crawled, and indexed for a URL?

Found means that a search engine knows the URL without necessarily having retrieved it. Crawled confirms a retrieval, not permanent processing; indexed means that a version can be considered for search results without guaranteeing a specific ranking.

Case test: "Crawl equals index"

A URL is listed in the sitemap and is therefore considered found, but it does not appear in any server log. A second URL was successfully retrieved, but processed as a duplicate of another canonical page; these two cases require completely different actions.

Source of evidence

Test criterion

Source of evidence

Sitemap, URL check, server log, and search result answer different status questions and are not equivalent.

Test criterion

Time reference

Each finding is assigned a timestamp because retrieval, processing, and index status can change independently.

  • Desired version – In the case of duplicates, it is necessary to check which canonical URL was processed, instead of only considering the entered address.

Desired version

  • Number and transition time from found to crawled and from crawled to indexed URLs for each class.

  • Percentage of checked URLs where the desired address is selected as the canonical indexed version.

Time reference

  1. The business query is first assigned to a state: knowledge, retrieval, processing, or actual delivery.

  2. URL checks, live tests, and server logs are compared side-by-side with timestamps and the desired canonical URL.

  3. Actions are then taken based on the earliest demonstrably missing step, rather than the vague term "indexing."

Crawl equals index

  • Crawl equals index A successful server response does not prove that content was selected or that the checked URL was accepted as canonical.

  • Index equals ranking Inclusion only creates the possibility of display and does not guarantee search query or ranking.

  • Outdated report Tools can show different processing times and provide seemingly contradictory snapshots.

How "Found, crawled, and indexed" is related to other decisions

Crawl Budget: When it's truly relevant and when it isn't Expands on the "Source of Evidence" checkpoint. The key question is: For which websites is crawl budget a real problem, and when is it merely a distraction?

A complementary perspective is offered Detecting Sitemap Errors Even When the File Is Formally ValidIt answers the question: "Which sitemap errors remain undetected despite a formally valid XML file?"

If you want to practically implement "distinguishing between found, crawled, and indexed," you can refer to Robust Website Systems This focuses on "index control and diagnosis" and "source of evidence."

Conclusion: Distinguishing between found, crawled, and indexed

These three states prevent premature diagnoses when they are documented with a source and time. Each stage has its own causes and its own appropriate checks.

Sources and Further Information

The classification of "found, crawled, and indexed" is based on the following official documentation and standards.

Key Thesis

Found means that the URL is known; crawled confirms a retrieval; indexed indicates inclusion as a search candidate. Links, logs, and index reports each verify a different step.

What This Is Not About

Found, crawled, and indexed are not interchangeable terms for visible or invisible in search.

What it's about

These states describe related but separate steps from URL recognition through retrieval to potential inclusion in a search index.

More insights

Crawling, Indexing & Canonicals

What "Crawled – not currently indexed" can actually mean

"Finding, crawled, and indexed" includes, as a separate verification step, the question: What are the actual possible causes of the status "Crawled, not currently indexed"?

Crawling, Indexing & Canonicals

Noindex pages in XML sitemaps: Why this is contradictory

Adds a separate decision to "Distinguishing between found, crawled, and indexed": Why shouldn't pages with noindex be included in the same XML sitemap?

Insights Overview

All VELUNO Insights at a Glance

Further analyses on Website Systems, digital visibility, and robust working models.

Practical Implications

Time reference: Path to approval

A specific problem case should initially be classified only as found, crawled, or indexed. Then, the exact transition at which the documented process breaks down is examined.