Langprotect

Common Crawl

A publicly available web dataset containing billions of webpages that is widely used to train and evaluate large language models.

What is Common Crawl?

Common Crawl regularly crawls billions of web pages and makes the collected data available through an open repository. Its datasets contain web page content, metadata, links, and other information gathered from across the public internet. Because of its scale, Common Crawl data has been widely used as a source for creating datasets used in natural language processing and training large language models.

Why is Common Crawl Important?

Training and researching modern AI systems often requires enormous amounts of text data. Common Crawl provides researchers and developers with access to large web datasets without requiring them to crawl the internet independently. However, using web-scale data also requires careful consideration of data quality, copyright, privacy, bias, and harmful content.

Common use cases

Common Crawl is commonly used for AI training datasets, natural language processing, web research, search technologies, data analysis, and large-scale language modeling.