OPUS 4 | Quantitative Linguistik

Weniger ist mehr? Eine Analyse zur „Neigung zum Hinzufügen“ im Deutschen anhand des neuen Häufigkeitsdatensatzes DeReKoGram (2024)

Wolfer, Sascha ; Koplenig, Alexander ; Kupietz, Marc ; Müller-Spitzer, Carolin

Languages with more speakers tend to be harder to (machine-)learn (2023)

Koplenig, Alexander ; Wolfer, Sascha

Computational language models (LMs), most notably exemplified by the widespread success of OpenAI's ChatGPT chatbot, show impressive performance on a wide range of linguistic tasks, thus providing cognitive science and linguistics with a computational working model to empirically study different aspects of human language. Here, we use LMs to test the hypothesis that languages with more speakers tend to be easier to learn. In two experiments, we train several LMs—ranging from very simple n-gram models to state-of-the-art deep neural networks—on written cross-linguistic corpus data covering 1293 different languages and statistically estimate learning difficulty. Using a variety of quantitative methods and machine learning techniques to account for phylogenetic relatedness and geographical proximity of languages, we show that there is robust evidence for a relationship between learning difficulty and speaker population size. However, contrary to expectations derived from previous research, our results suggest that languages with more speakers tend to be harder to learn.

A large quantitative analysis of written language challenges the idea that all languages are equally complex (2023)

Koplenig, Alexander ; Wolfer, Sascha ; Meyer, Peter

One of the fundamental questions about human language is whether all languages are equally complex. Here, we approach this question from an information-theoretic perspective. We present a large scale quantitative cross-linguistic analysis of written language by training a language model on more than 6500 different documents as represented in 41 multilingual text collections consisting of ~ 3.5 billion words or ~ 9.0 billion characters and covering 2069 different languages that are spoken as a native language by more than 90% of the world population. We statistically infer the entropy of each language model as an index of what we call average prediction complexity. We compare complexity rankings across corpora and show that a language that tends to be more complex than another language in one corpus also tends to be more complex in another corpus. In addition, we show that speaker population size predicts entropy. We argue that both results constitute evidence against the equi-complexity hypothesis from an information-theoretic perspective.

cOWIDplus Viewer: Sprachliche Spuren der Corona-Krise in deutschen Online-Nachrichtenmeldungen. Explorieren Sie selbst! (2020)

Müller-Spitzer, Carolin ; Wolfer, Sascha ; Koplenig, Alexander ; Michaelis, Frank

Studying Lexical Dynamics and Language Change via Generalized Entropies: The Problem of Sample Size (2020)

Koplenig, Alexander ; Wolfer, Sascha ; Müller-Spitzer, Carolin

Recently, it was demonstrated that generalized entropies of order α offer novel and important opportunities to quantify the similarity of symbol sequences where α is a free parameter. Varying this parameter makes it possible to magnify differences between different texts at specific scales of the corresponding word frequency spectrum. For the analysis of the statistical properties of natural languages, this is especially interesting, because textual data are characterized by Zipf’s law, i.e., there are very few word types that occur very often (e.g., function words expressing grammatical relationships) and many word types with a very low frequency (e.g., content words carrying most of the meaning of a sentence). Here, this approach is systematically and empirically studied by analyzing the lexical dynamics of the German weekly news magazine Der Spiegel (consisting of approximately 365,000 articles and 237,000,000 words that were published between 1947 and 2017). We show that, analogous to most other measures in quantitative linguistics, similarity measures based on generalized entropies depend heavily on the sample size (i.e., text length). We argue that this makes it difficult to quantify lexical dynamics and language change and show that standard sampling approaches do not solve this problem. We discuss the consequences of the results for the statistical analysis of languages.

cOWIDplus Analyse: Wie sehr schränkt die Corona-Krise das Vokabular deutschsprachiger Online-Presse ein? (2020)

Wolfer, Sascha ; Koplenig, Alexander ; Michaelis, Frank ; Müller-Spitzer, Carolin

cOWIDplus Analyse ist eine kontinuierlich aktualisierte Ressource zu der Frage, ob und wie stark sich der Wortschatz ausgewählter deutscher Online-Pressemeldungen während der Corona-Pandemie systematisch einschränkt und ob bzw. wann sich das Vokabular nach der Krise wieder ausweitet. In diesem Artikel erläutern die Autor*innen die hinter der Ressource stehende Forschungsfrage, die zugrunde gelegten Daten, die Methode sowie die bisherigen Ergebnisse.

cOWIDplus Viewer: Sprachliche Spuren der Corona-Krise in deutschen Online-Nachrichtenmeldungen. Explorieren Sie selbst! (2020)

Müller-Spitzer, Carolin ; Wolfer, Sascha ; Koplenig, Alexander ; Michaelis, Frank

cOWIDplus Viewer (2020)

Wolfer, Sascha ; Koplenig, Alexander ; Michaelis, Frank ; Müller-Spitzer, Carolin

cOWIDplus (2020)

Wolfer, Sascha ; Koplenig, Alexander ; Michaelis, Frank ; Müller-Spitzer, Carolin

Die Corona-Krise hat Einfluss auf die Sprache in deutschsprachigen Online-Medien. Wir haben die Hypothese, dass sich die Vielfältigkeit des verwendeten Vokabulars einschränkt. Wir glauben zudem, dass sich die Diversität des Vokabulars nach "überstandener" Krise wieder auf ein "Prä-Pandemie-Niveau" einpendeln wird. Diese zweite Hypothese lässt sich erst im Laufe der Zeit überprüfen.

Studying Lexical Dynamics and Language Change via Generalized Entropies: The Problem of Sample Size (2019)

Koplenig, Alexander ; Wolfer, Sascha ; Müller-Spitzer, Carolin

Recently, it was demonstrated that generalized entropies of order α offer novel and important opportunities to quantify the similarity of symbol sequences where α is a free parameter. Varying this parameter makes it possible to magnify differences between different texts at specific scales of the corresponding word frequency spectrum. For the analysis of the statistical properties of natural languages, this is especially interesting, because textual data are characterized by Zipf’s law, i.e., there are very few word types that occur very often (e.g., function words expressing grammatical relationships) and many word types with a very low frequency (e.g., content words carrying most of the meaning of a sentence). Here, this approach is systematically and empirically studied by analyzing the lexical dynamics of the German weekly news magazine Der Spiegel (consisting of approximately 365,000 articles and 237,000,000 words that were published between 1947 and 2017). We show that, analogous to most other measures in quantitative linguistics, similarity measures based on generalized entropies depend heavily on the sample size (i.e., text length). We argue that this makes it difficult to quantify lexical dynamics and language change and show that standard sampling approaches do not solve this problem. We discuss the consequences of the results for the statistical analysis of languages.

Open Access

Quantitative Linguistik

Refine

Author

Year of publication

Document Type

Language

Has Fulltext

Is part of the Bibliography

Keywords

Publicationstate

Reviewstate

Publisher

14 search hits