Data & Databases
Common Crawl
A nonprofit project that publishes collections of information derived from web crawling.
Example
Researchers use its collected web pages as material for further study.
Why people use it
It makes large web collections available for research and other data work.
What you'll hear
“Does this collection include material from Common Crawl?”
What this means for you
Check the pages' origins, quality and permitted uses before relying on them.
Can you control it?
No
No direct control. This describes a wider issue, concept or result rather than something you can simply switch on or off in a tool.
Common questions
- Does inclusion in Common Crawl settle all reuse rights or quality questions?
- No. Public collection of information access does not remove basic rights and data-quality considerations.
- Does it contain the entire internet?
- No. Web collections have gaps and reflect what was reachable when the crawl occurred.
- Is every page in a collection current?
- No. A stored page reflects a past visit and may differ from the website today.