Web crawler
internet bot that systematically browses the World Wide Web, typically for the purpose of Web indexing (web spidering)

A web crawler, sometimes called a spider or spiderbot and often shortened to crawler, is an Internet bot that systematically browses the World Wide Web and that is typically operated by search engines for the purpose of Web indexing (web spidering).
Web search engines and some other websites use web crawling or spidering software to update their web content or indices of other sites' web content. Web crawlers copy pages for processing by a search engine, which indexes the downloaded pages so that users can search more efficiently.
The number of web pages is extremely large, and search engines do not index all web content. Studies of late-1990s search engines found that individual engines indexed only a fraction of the then-indexable web. Modern search engines use crawling, indexing, and ranking systems to return relevant results quickly, although not all pages are crawled, indexed, or served.
Crawlers can validate hyperlinks and HTML code. They can also be used for web scraping and data-driven programming.
Nomenclature
A web crawler is also known as a spider, an ant, an automatic indexer, or (in the FOAF software context) a Web scutter.
Overview
A web crawler starts with a list of URLs to visit.
“Web crawler” enters the record as internet bot that systematically browses the World Wide Web, typically for the purpose of Web indexing (web spidering). Crown Archives preserves that source wording while asking what crawler, internet and systematically can confirm, complicate or overturn.
Why this record matters
“Web crawler” is worth following because a concise public description often conceals a longer documentary argument. Here, crawler, internet and systematically provides the most credible route into that argument.
Named sources, stable identifiers and responsible institutions provide the strongest route from overview to verifiable evidence. The source revision retrieved here is dated Sep 20, 2026. The linked authority identifier is Q45842. None of the 0 selected statements returned an explicit reference.
A concise general-reference account can conceal disagreements about scope, terminology or the weight assigned to individual sources. The source lead contains qualifying language; that uncertainty should survive quotation, summary and reuse. Authority statements aid reconciliation but still require their own references, qualifiers and ranks to be checked.
How to read it
Use the entry as an orientation point, then follow its citations and revision history. Names, dates and institutional relationships should be checked against the original record.
- Subject orientation
- Search vocabulary
- Locating named sources
The closest primary source, responsible institution and strongest cited specialist reference.
Three-step research path
- Establish the record: confirm the title “Web crawler”, its source revision and the description used here.
- Expand the search: follow Web crawler primary sources, Web crawler archive and crawler research across catalogues and specialist indexes.
- Test the account: compare the strongest cited source with the responsible institution’s current record and note any disagreement.
Questions for further research
- Which source most directly establishes the central claim about “Web crawler”?
- What terminology or title could unlock a more precise catalogue search?
- Which institution is responsible for the underlying evidence?
Search terms from this dossier
This entry incorporates text from “Web crawler” on English Wikipedia. Contributors are listed in the page history. Text is available under the Creative Commons Attribution-ShareAlike 4.0 License. Selected authority identifiers and statements are retrieved from Wikidata under CC0; their references and qualifiers remain part of the verification path.