Volltext-Downloads (blau) und Frontdoor-Views (grau)

Ungoliant: An optimized pipeline for the generation of a very large-scale multilingual web corpus

  • Since the introduction of large language models in Natural Language Processing, large raw corpora have played a crucial role in Computational Linguistics. However, most of these large raw corpora are either available only for English or not available to the general public due to copyright issues. Nevertheless, there are some examples of freely available multilingual corpora for training Deep Learning NLP models, such as the OSCAR and Paracrawl corpora. However, they have quality issues, especially for low-resource languages. Moreover, recreating or updating these corpora is very complex. In this work, we try to reproduce and improve the goclassy pipeline used to create the OSCAR corpus. We propose a new pipeline that is faster, modular, parameterizable, and well documented. We use it to create a corpus similar to OSCAR but larger and based on recent data. Also, unlike OSCAR, the metadata information is at the document level. We release our pipeline under an open source license and publish the corpus under a research-only license.

Download full text files

Export metadata

Additional Services

Search Google Scholar

Statistics

frontdoor_oas
Metadaten
Author:Julien Abadji, Pedro Javier Ortiz SuárezORCiD, Laurent RomaryORCiDGND, Benoît SagotORCiD
URN:urn:nbn:de:bsz:mh39-104688
DOI:https://doi.org/10.14618/ids-pub-10468
Parent Title (English):Proceedings of the Workshop on Challenges in the Management of Large Corpora (CMLC-9) 2021. Limerick, 12 July 2021 (Online-Event)
Publisher:Leibniz-Institut für Deutsche Sprache
Place of publication:Mannheim
Editor:Harald Lüngen, Marc Kupietz, Piotr Bański, Adrien Barbaresi, Simon Clematide, Ines Pisetta
Document Type:Conference Proceeding
Language:English
Year of first Publication:2021
Date of Publication (online):2021/06/23
Publicationstate:Veröffentlichungsversion
Reviewstate:Peer-Review
Tag:corpus linguistics; large corpora
GND Keyword:Automatische Sprachanalyse; Computerlinguistik; Korpus <Linguistik>; Natürliche Sprache; Open Source; Urheberrecht
First Page:1
Last Page:9
DDC classes:400 Sprache / 400 Sprache, Linguistik
Open Access?:ja
Linguistics-Classification:Korpuslinguistik
Conferences, Workshops:Proceedings of the Workshop on Challenges in the Management of Large Corpora (CMLC-9) 2021. Limerick, 12 July 2021 (Online-Event)
Licence (German):License LogoCreative Commons - CC BY - Namensnennung 4.0 International