Common Crawl maintains a free,open repository of web crawl data that can be used by anyone.
Common Crawl is a 501(c)(3) nonâprofit founded in 2007. â We make wholesale extraction, transformation and analysis of open web data accessible to researchers.
We are pleased to announce that the crawl archive for August 2026 is now available, containing 2.14 billion web pages or 360 TiB of uncompressed content.