Skip to content

Data & Databases

Common Crawl

A nonprofit project that publishes collections of information derived from web crawling.

Example

Researchers use its collected web pages as material for further study.

Why people use it

It makes large web collections available for research and other data work.

What you'll hear

“Does this collection include material from Common Crawl?”

What this means for you

Check the pages' origins, quality and permitted uses before relying on them.

Can you control it?

No

No direct control. This describes a wider issue, concept or result rather than something you can simply switch on or off in a tool.

Common questions

Does inclusion in Common Crawl settle all reuse rights or quality questions?
No. Public collection of information access does not remove basic rights and data-quality considerations.
Does it contain the entire internet?
No. Web collections have gaps and reflect what was reachable when the crawl occurred.
Is every page in a collection current?
No. A stored page reflects a past visit and may differ from the website today.

Related terms

Still have questions?

Up to 500 characters.

Ask LATHIC about AI. Relevant glossary entries may be included.

Your question, the glossary entries it matches, and a rotating pseudonymous identifier go to Microsoft Azure’s OpenAI service through Vercel AI Gateway to generate an answer. Zero retention and no training are required of the provider, and LATHIC does not save your question or answer. Privacy Notice