OPUS 4 | Search

5 search hits

1 to 5

Sort by

Automatic classification of Russian texts for didactic purposes (2017)

Batinić, Dolores ; Birzer, Sandra ; Zinsmeister, Heike

In this paper we present the results of an automatic classification of Russian texts into three levels of difficulty. Our aim is to build a study corpus of Russian, in which a L2 student is able to select texts of a desired complexity. We are building on a pilot study, in which we classified Russian texts into two levels of difficulty. In the current paper, we apply the classification to an extended corpus of 577 labelled texts. The best-performing combination of features achieves an accuracy of 0,74 within at most one level difference.

Corpus Masking: Legally Bypassing Licensing Restrictions for the Free Distribution of Text Collections (2007)

Rehm, Georg ; Witt, Andreas ; Zinsmeister, Heike ; Dellert, Johannes

Creating an extensible, levelled study corpus of Russian (2016)

Batinić, Dolores ; Birzer, Sandra ; Zinsmeister, Heike

In this paper, we present first results of training a classifier for discriminating Russian texts into different levels of difficulty. For the classification we considered both surface-oriented features adopted from readability assessments and more linguistically informed, positional features to classify texts into two levels of difficulty. This text classification is the main focus of our Levelled Study Corpus of Russian (LeStCoR), in which we aim to build a corpus adapted for language learning purposes – selecting simpler texts for beginner second language learners and more complex texts for advanced learners. The most discriminative feature in our pilot study was a lexical feature that approximates accessibility of the vocabulary by the second language learner in terms of the proportion of familiar words in the texts. The best feature setting achieved an accuracy of 0.91 on a pilot corpus of 209 texts.

Digital Text Resources for the Humanities – Legal Issues (2007)

Rehm, Georg ; Witt, Andreas ; Hinrichs, Erhard ; Lehmberg, Timm ; Chiarcos, Christian ; Zimmermann, Felix ; Zinsmeister, Heike ; Dellert, Johannes

Linguistically Annotated Corpora: Quality Assurance, Reusability and Sustainability (2008)

Zinsmeister, Heike ; Witt, Andreas ; Kübler, Sandra ; Hinrichs, Erhard